But it wasn't explicitly told to hack HuggingFace. It was told "answer this security question", and it's answer was to break into the teacher's desk to find the answer key.
HN user
gallerdude
If the models were the quality of GPT-5.6, or Opus 4.8, would Meta really reap any benefits? Maybe in a zero-sum way, because OpenAI and Anthropic would lose. But I'm not sure they'd be that much more ahead either.
Staying at home is as entertaining as it’s ever been: video games, Netflix, don’t even get me started on short-form content.
The original comment was “open models are what kill OpenAI and Anthropic”, which to me is as silly as saying “Android is what killed Apple”
Android has market share, but Apple makes all of the money! I find it really funny when people attribute Apple’s success to “oh, the only reason they succeed is design and marketing.” Yeah, I mean factually speaking design and marketing actually do matter a lot!
Us developer types like to pretend like specs are the only thing that matters? If you could have a 10x more powerful model you could only access running locally through your terminal, versus a weaker model through a clean web interface, normies will pick the web ui every single time. Product experience is simply everything, as much as we like to pretend like nitty technical decisions are the most important thing.
Even in the world where all models are basically equivalent (a thesis I don’t buy, but will grant you for arguments sake) - I believe there is much more to the AI business than just training and running models.
It’s a very new set of technologies, and understanding what is useful to customers and what isn’t is the whole game. Call it, product taste. There were a million cell phones before the iPhone took over the world. Why iPhone? Product taste. There are a million startups, and only a select few become unicorns. Why? Product taste.
This was a really fun visualization, so I vibecoded it.
Why would you use AI to write this post? If you can’t bother to write it, why should I bother to read it?
I think the question “would China cooperate” needs much more investigation. Everyone online pundit seems to think “obviously not”, but they’re people too with clear positive and negative incentives. It’s possible they’ve found a very similar calculus that we have.
“Politics is the art of the possible”
If you’ve done any software development at all, certainly.
Of course. But is it really impossible that Dario’s directive to the marketing team is “try not to make us look bad, but also be honest about our models’ capabilities, so people can stay informed”?
I think they're pretty happy to have government sharing some of the big societal responsibility, honestly.
Well it seems both Anthropic and OpenAI are consciously choosing to do this, which means, for now, neither plan on suing. So if no one sues, how could it be unconstitutional?
So let’s say you’re in Anthropic’s shoes. You see that LLM’s are getting better and better, and it’s very possible that they will have some impact on jobs in the next few years, and a very meaningful impact on cybersecurity.
Is it more ethical to stay silent about these concerns, as you might have a bit of self interest? Or even if it looks a bit self interested, is it better to warn people ahead of time? I think the latter is obviously the better position.
I really liked Dario's metaphor that in the 80's, we could have said someday we'll have "supercomputers", which can do all the calculations we did except WAY faster. When, in reality, the AI's just get smarter over time, even if the frontier is jagged. AGI is just vibes only for "smart enough, consistently enough".
I was a baby when the Internet Revolution happened. I was in high school and college when the Mobile Revolution steamrolled everything. It’s been interesting to see this one, as an adult working in the world. I wonder how far it will go.
I actually have no problem with the 5.x line... but if Pro really was an entirely new pretrain, they did a horrible job conveying that.
If GPT-5.5 Pro really was Spud, and two years of pretraining culminated in one release, WOW, you cannot feel it at all from this announcement. If OpenAI wants to know why they like they’ve fallen behind the vibes of Anthropic, they need to look no further than their marketing department. This makes everything feel like a completely linear upgrade in every way.
I agree partially, but also misses the wonder he would have for: relaxing bathtubs, funny livestreams, wireless earbuds, huge libraries, and even globes.
And yeah, you could make a list of struggles we have today he never did. But that’s kind of my point - it’s complicated.
1. You can't understand the nuances, but there is a general pattern: new inventions may make us slightly less proficient at specifics, yet more powerful overall
2. Imagine a hunter gatherer is time travelled to 2026. You have lunch go to a cafe with him, and he learns that food is cheap, delicious, and abundant. He sees your house, and thinks it's amazing compared to his cave. He thinks that 2026 must be absolute paradise. You explain to him, well kinda, but also not really. Is the hunter gatherer right?
I’ve been thinking about AI robotics lately… if internally at labs they have a GPT-2, GPT-3 “equivalent” for robotics, you can’t really release that. If a robot unloading your dishwasher breaks one of your dishes once, this is a massive failure.
So there might be awesome progress behind the scenes, just not ready for the general public.
But rolling your own can’t be that much cheaper than buying it from a leading lab. Especially when you consider the amount of spending on datacenters.
I’m sure there’s more to it than this, but it feels like Zuck has pet interests like VR and now AI.
I’m not sure. If it was open source, certainly. But 4th place doesn’t really matter if you have nothing different to add.
This would have been an amazing release 6 months ago. But the industry moves so fast, this is a trite release. Maybe it’s best for Meta to sell their superintelligence division. I don’t think Zuck’s vision is particularly compelling.
Tbf there was a 5.3 codex
My job may have become part of the training data with how much coverage there is around it. Perhaps another career would be a better test of LLM capabilities.
We used to get one annual release which was 2x as good, now we get quarterly releases which are 25% better. So annually, we’re now at 2.4x better.
The weirdest thing about this AI revolution is how smooth and continuous it is. If you look closely at differences between 4.6 and 4.5, it’s hard to see the subtle details.
A year ago today, Sonnet 3.5 (new), was the newest model. A week later, Sonnet 3.7 would be released.
Even 3.7 feels like ancient history! But in the gradient of 3.5 to 3.5 (new) to 3.7 to 4 to 4.1 to 4.5, I can’t think of one moment where I saw everything change. Even with all the noise in the headlines, it’s still been a silent revolution.
Am I just a believer in an emperor with no clothes? Or, somehow, against all probability and plausibility, are we all still early?
I always grew up hearing “competition is good for the consumer.” But I never really internalized how good fierce battles for market share are. The amount of competition in a space is directly proportional to how good the results are for consumers.