I suspect that distillation attacks may be slightly exaggerated. Most of the training data used during fine-tuning is now synthetic data. You can't just repeat the same stuff twice, therefore another LLM is writing a text book that is explaining a topic in detail, ideally without any gaps in the material.
HN user
MichaelMoser123
Actually here is the full talk on youtube, the quote is at 20:26 https://www.youtube.com/watch?v=t9HmOz8H0qI
I saw a video on Linked-in. It's a talk by Steve Jobs from 1983, where he is anticipating, what an LLM can do now (at 4:03 in the video). https://lnkd.in/p/ddWU7cTp
My thoughts on the subject: What would Aristotle have said? The LLM is simulating Aristotle's response, based on syntactic and pragmatic dependencies between words and phrases taken from his works. Often you will get a sensible answers, but sometimes you won't. It will require some understanding of the larger context to tell the difference, you again need to read some books in order to tell the difference and in order to engage in a meaningful Socratic dialogue. Without human scrutiny you will get wrong results, and you need some real training to handle that.
just recalled: some significant open source projects are being rewritten in Rust, with the help of Claude. I don't know if these efforts will be successful in the long run, but in some way these news headlines may be creating a media dynamic as part of the IPO preparations?
[1] https://news.ycombinator.com/item?id=48870966 pgrust passes 100% of the Postgres regression tests
[2] https://news.ycombinator.com/item?id=48837877 Rewriting Bun in Rust
[3] https://news.ycombinator.com/item?id=48789325 My AI-built PHP engine in Rust passes 17% of PHP-src tests, renders WordPress (ekinertac.com)
Speaking of financing: how is the Anthropic IPO going, what is the timeline? They filed over a month ago, no news since. (I would have expected some spectacular news headlines that would be designed to fuel public interest in the impending IPO, but can't detect anything of substance)
Wow. Now did you try to check the setup with something like Claude Fable? Will it find issues, what kind of issues? Another question: how many tokens did this effort cost? Did you learn new prompting tricks?
i happened to liked Google AI mode, even wrote a composition about that [1] Now it is going to be enshittified, which is probably inevitable.
https://github.com/MoserMichael/tips_on_using_google_ai_mode
CNN is quoting data from the Gaza health ministry, an organization run by genocidal Islamist Hamas, without mentioning its affiliation and without questioning the data. So much for objectivity on CNN and "Hacker News". There is also no mention on the food convoys that get plundered by Hamas. Just to mention: Hamas is designated as a terrorist organization in the US, since 1997 - just adding some missing background information.
I think publishers will not be pleased about a steep fall in click rates, now publishers still have considerable political influence. What will Google do in the event of serious legal pushback or a renewed drive for antitrust action?
The introduction of AI overviews into Google search will cost quite a lot in compute/other resources, despite heavy caching, therefore this might be a significant bet in terms of costs vs profit for Google. What does Google expect from this feature in terms of business results? This seems to be quite a big bet, but what is actually at stake - in real terms?
Come to think of it: is there now a showdown between Google and Microsoft/OpenAI, where collateral costs are no longer taken into account?
I seriously think I'm going to stop posting on the internet for good.
I had similar thoughts, but it would probably not make a difference, at this stage. What is there stays there - either online, as in the case of HN, or as part of some collected dataset.
In hindsight: the world changed in so many ways, from the world I knew some twenty years ago, and I am not even talking about politics or technology: the attitudes and perception of people seems to have changed in many ways. Back then I thought it would be of benefit to be open and upfront about things. Now that is no longer a common perception.
Enough said.
reliability is much better now, as far as i can tell.
using zeebe/Camunda at work. The system gives you a way of designing and partitioning message-based workflows. It has a very thorough design.
"What do you think about user XYZ?" or "What do you think about the comments of user XYZ?"
It starts a whole lot of SQL queries that find and aggregate data & statistics
It must have a very interesting and well written system prompt for this type of questions.
(gives me second thoughts about my personal approach to privacy)
I once had a python side project, it parses the 1911 edition of Roget Thesaurus into memory and provides some queries.
cpython doesn't have a JIT, why is free-threaded python a higher priority than developing a just in time compiler? The later would be more resonant with the typical use case for python and benefit a larger portion of users, wouldn't it? (Wouldn't a backend server project use golang or java to begin with?)
you have a point. The usual approach was to choose a subset of C++ features, so it becomes appropriate for a given project. But yes, the language is huge - as it tries to suite everyone but no one in particular (which is insane)
and putting structure instances into an array so that you can refer to them via indexes of the array entries (as the only escape from being maimed by the borrow checker) is normal?
deepseek-v2,v3,r1 are all using multi-headed attention.
I hope you are right, however:
https://en.wikipedia.org/wiki/Indus_Waters_Treaty#Suspension
Following the suspension of the treaty, India significantly reduced the flow of water through the Chenab River, which crosses into Pakistan. Pakistani authorities claimed a 90% drop in water supply and accused India of choking the river’s flow. India also initiated new hydroelectric projects and began constructing dams on the western rivers, actions previously constrained under the treaty.[125][126][127]
Pakistan has reportedly warned that any attempt by India to disrupt the flow of water from shared rivers could be considered an act of war, and would attack India with nuclear weapons.[128]
The bad news: there is some real potential for escalation due to the suspension of the Indus Waters Treaty
https://economictimes.indiatimes.com/news/india/indias-water...
https://en.wikipedia.org/wiki/Indus_Waters_Treaty
Wasn't there something in the intro of "Mad Max fury road" about water wars?
i think an Agent is an LLM that interacts with the outside world via a protocol like MCP, that's a kind of REST-like protocol with a detailed description for the LLM on how to use it. An example is an MCP server that knows how to look up the price for a given stock ticker, so it enables the LLM to tell the current price for that ticker.
see: https://github.com/luigiajah/mcp-stocks
The implementation: https://github.com/luigiajah/mcp-stocks/blob/main/main.py
Each MCP endpoint comes with a detailed comment - that comment will be part of the metadata published by the MCP server / extension. The LLM reads this instruction when the MCP extension is added by the end user, so it will know how to call it.
The main difference between REST an MCP is that MCP can maintain state for the current session (that's an option), while REST is supposed to be inherently stateless.
I think most of the other protocols are a variation of MCP.
Are all of the ai model benchmarks just made up? https://epoch.ai/data/ai-benchmarking-dashboard https://ai.azure.com/explore/benchmarks
At the time of the Pentium people could work as graphic artists, without competition from generative AI models. Education was the differentiator, that used to be a settled question.
We noticed the sincerity of this outreach on the 7th of October 2023, with 1,195 people butchered by the terrorists and 251 hostages taken. We are still grappling with the results.
Thanks for your comment, it explains why there seems to be a high degree of support for these measures in some quarters (was looking at the youtube comments to the liberation day speech) vs the consensus here at HN.
Let's suppose these policies are to the benefit of some Americans over other the benefit of other Americans. The open question now is: does it matter? does it really have an influence on the gross profit numbers? Will an isolationist foreign policy destroy the international order and how could this effect the US in return?
Naive question of a foreigner: I see a lot of approval on the youtube comments to Trump's liberation day speech and over at twitter. Now the consensus on HN is diametrically opposed. My question: why do some people think that this is a good idea? Why is there such a huge rift in US society and how do you plan to bridge this gap?
In Soviet Russia AI Dev Tools organize you
(couldn't resist the urge to post slashdot-like silliness)
yes, however if you look at the function definition then you can't tell if the caller is passing a slice vs an array, or an interface that is implemented on a pointer vs a struct.
The semantics of the function parameter should not depend on how the function is being called.
Another thing that bugs me: if you have a function argument then it is copied by value, right? But it is hard to tell if the whole value is copied or not.
- structures argument: the whole structure is copied (unless you passed a pointer to the struct)
- interface? depends if the implementation is structure or pointer.
- maps: it copies just a small internal struct that is pointing to the implementation of the map.
- arrays: depends if it is a slice then copy cost is small (again, similar to maps), however an array is fully copied.
... complicated.
"Culture fit" is pretty much trading off communication efficiency and high trust against diversity of thought.
I think I understand now: if the structure of your company is strictly top down, then you will have to value "communication efficiency" and "high trust" criteria higher than "diversity of thought".
However, you might not be able to pivot efficiently, if your core assumptions are disproved. That's the situation where you might need "diversity of thought" - and the ability to incorporate different kinds of feedback.
Though I don't quite think that you will find this insight in this frigging book.