A lot of us went through a similar system and still retained both honour and desire to actually learn.
HN user
jmmcd
I'm a lecturer and researcher in computer science and AI in Ireland. I'm into music, computer music, programs, and programs that write programs.
It's interesting that Claude also over-uses en-dashes. It's very willing to create compound-noun-phrases, especially in that compressed-summary-paragraph it often writes. The 0-days-vibes-vulns that started this thread looks a lot like that, but it could be Claude directly, or just Claude's style influencing people who spend too much time with it.
People are missing that Willison is among the very best people we have in the role of (for lack of a good name): early access to frontier models, evaluate them in real scenarios, no wishful thinking, hype, or doom, communicate the possibilities. Yes he could have fixed this himself but then he would have learned nothing about the AI, and we wouldn't have read a fascinating and important article.
LLMs do just interpolate their training data
"interpolate" has a technical meaning - in this meaning, LLMs almost never interpolate. It also has a very vague everyday meaning - in this meaning, LLMs do interpolate, but so do humans.
But euros spent on tokens is a tiny fraction of the overall costs of the project.
Since these companies can’t improve their AI models without fresh data created by human beings
Totally wrong. Self-play dates back to Arthur Samuel in the 1950s and RL with verifiable rewards is a key part of training the most advanced models today.
Anglo
Please, write US-American. These people are not coming from any other place.
I understand your point, but in response to GP (they should spend this money on houses for other poor people instead), the reduced reliance on other social welfare is totally legitimate to count.
You didn't read the article. The scheme gave positive return on investment.
Google Scholar provides imperfect citations - very often wrong article type (eg article versus conference paper), but up to and including missing authors, in my experience.
Jokes are one of the good parts of human existence, so - while I see your side of the story - there is another side.
The best example of all is Prolog. It is always held up as the paradigmatic representative of logic programming, a rare language paradigm. But it doesn't need to be a language. It is really a collection of algorithms which should be a library in every language, together with a nice convention for expressing Prolog things in that language's syntax.
(My comment is slightly off-topic to the article but on-topic to the title.)
But he actually uses frontier LLMs in his own work. Probably that's stronger evidence.
Yes. But "the gold standard" just means "the most natural, easy and dumb way".
This is on the semi-private set
"Pelican on bicycle" is one special case, but the problem (and the interesting point) is that with LLMs, they are always generalising. If a lab focussed specially on pelicans on bicycles, they would as a by-product improve performance on, say, tigers on rollercoasters. This is new and counter-intuitive to most ML/AI people.
About ARC 2:
I would want to hear more detail about prompts, frameworks, thinking time, etc., but they don't matter too much. The main caveat would be that this is probably on the public test set, so could be in pretraining, and there could even be some ARC-focussed post-training - I think we don't know yet and might never know.
But for any reasonable setup, if no egregious cheating, that is an amazing score on ARC 2.
On HN it's very common to see a blog post along the lines of "I found this old piece of equipment with no brand name, I used some network traffic inspection to figure out what it does, I hacked around a bit, I got it working and turned it into a self-ringing doorbell with wifi" (or whatever). All of that is anecdotal, N=1, "I did what worked for me, I hope it's interesting to you". And those posts are highly prized and rightly so.
(a) no it's not
(b) your comment is miles off-topic, as he is not addressing doom in any sense
But they're submissions to ICML.
It could certainly replace the author of this article.
Modern LLMs can one-shot code in a totally new language, if you provide the language manual. And you have to provide the language manual, because otherwise how can the students learn the language.
The fact that AI can do your homework should tell you how much your homework is worth.
A lot of people who say this kind of thing have, frankly, a very shallow view of what homework is. A lot of homework can be easily done by AI, or by a calculator, or by Wikipedia, or by looking up the textbook. That doesn't invalidate it as homework at all. We're trying to scaffold skills in your brain. It also didn't invalidate it as assessment in the past, because (eg) small kids don't have calculators, and (eg) kids who learn to look up the textbook are learning multiple skills in addition to the knowledge they're looking up. But things have changed now.
Absolutely devastating for the credibility of FAIR.
Link [1] doesn't seem to mention svg or vector graphics at all.
This important paper from Anthropic includes evidence that part (but only part) of reasoning is cross-lingual:
https://www.anthropic.com/research/tracing-thoughts-language...
No, you're confusing GA with GP.
in genetic programming the goal is to find/optimize dataset to fit given algorithm
No. Possibly you're confused between GAs and GP, a common confusion. In GP, the goal is to find an algorithm - in a real programming language, not as weights - to optimise an objective function. Often that is specified by input-output pairs, similar to supervised learning.
Yes indeed. This is because it's very disconcerting to see yourself non-mirrored.
"During videoconferencing, of course, you know how you look, since you can see yourself too," Walter-Terrill noted. "But on a call with dozens of people, you may be the only one who doesn't know how you sound to everyone else: you may hear yourself as rich and resonant, while everyone else hears a tinny voice."
Not quite! You don't really know how you look, because you see yourself before transmission.