But this incident undermines OpenAI's safety case compared with open models.
HN user
0xDEAFBEAD
Worried about irresponsible AI firms? Consider contributing to the campaign of Alex Bores, running for Congress in NY-12 https://ericneyman.wordpress.com/2025/10/20/consider-donating-to-alex-bores-author-of-the-raise-act/
You can try to contact me if you want, by emailing my username on protonmail, but I probably won't see it. Sorry.
An AI which is mistaken about the goal you give it, or how to go about achieving it, will behave in a de facto malicious manner, as this incident illustrates.
The AI can figure out whether it's airgapped. So its deployment behavior could be much different from the test behavior, when it's inevitably connected to the internet during deployment.
I mean, HuggingFace contacted law enforcement about this breach. That seems a little different to me.
Mythos established that these capabilities existed. This incident establishes that we can't control them.
Their models, very famously, are prone to reward hacking benchmarks in ways that other models are not. They need to publish numbers showing that their models are just as good as Anthropic's, since their entire business is at risk of collapsing if everyone is aware of how behind the frontier they truly are.
This doesn't seem internally consistent.
This incident basically announces to the world the message that "our models are prone to reward hacking". That renders any published benchmark numbers suspect. It also undermines the case for using OpenAI projects in business-critical applications--the exact application area where they might be able to sustain a moat against open-weight models.
There is a lot of conspiratorial thinking in this thread. I think people are engaging in wishful thinking to avoid cognitive dissonance from the possibility that we are in an increasingly dire situation. I would encourage people to sit with this possibility for a few minutes if they haven't already.
You're aware that HuggingFace notified law enforcement about this incident? Was that OpenAI's intended outcome when they prompted their AI?
Dario has more or less assented to an AI development pause
https://xcancel.com/AISafetyMemes/status/2014018200325722348...
I don't think we should be running cover for continued reckless AI development.
OpenAI already has loads of publicity. At this point, they don't need more brand recognition. This incident just has the effect of tarnishing their brand.
OpenAI leadership has been lobbying against regulation of AI systems. That doesn't comport with instigating incidents like this one, which give ammo to the heavy-regulation advocates.
"Models don't kill people. People kill people."
If that's the plan, today's failure by OpenAI looks really bad for any regulator who is trying to figure out whether to give OpenAI a license.
Any sort of warning or failure can always be written off as "marketing" to provide comfortable reassurance that there is no cause for alarm. There is an element of wishful thinking driving it, in my opinion.
What sort of warning or failure would be evidence against the "marketing" claims? Do we need to wait for a mass casualty event?
Best practice in safety engineering is to understand, diagnose, and respond to even small failures.
Why has Sam Altman worked to undermine doomers and downplay doom fears, if he benefits from incidents like this due to marketing?
https://xcancel.com/HumanHarlan/status/1965932275465597077#m
https://xcancel.com/AISafetyMemes/status/2062254769402699922...
It's an alignment problem in the sense that it demonstrates the principle that today's AI systems cannot be trusted to reliably work towards the goals of their users. A small-scale alignment failure and a large-scale alignment failure are the same fundamental type of failure. Typically, large disasters come after smaller disasters which foreshadowed the disaster mechanism, but weren't taken seriously.
Imagine if Hiroshima and Nagasaki were never destroyed. People would be arguing that fear of nuclear war is "just another religion".
Recall that the Morris Worm was designed as a harmless proof of concept, but ended up taking down 10% of the internet. Exponential growth can quickly get out of control. You would think that people would've learned that lesson from COVID.
I see a number of China-based signatories on this open letter signed by a bunch of luminaries
"Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war."
https://aistatement.com/work/statement-on-ai-extinction-risk...
You guys have created this un-falsifiable "marketing" narrative. Why is it that Jensen is pushing back on the doomer stuff, and complaining that it is hurting AI investments?
https://www.businessinsider.com/nvidia-jensen-huang-ai-doome...
Indeed. The Economist wrote an article a couple years ago arguing that Xi has been influenced by a Chinese Turing Award winner who believes AI poses a greater existential risk to humans than nuclear or biological weapons.
https://www.economist.com/china/2024/08/25/is-xi-jinping-an-...
Attacking a person instead of the substance of what they're saying is ad hominem.
Hasn't the AI 2027 "creative writing exercise" held up not-so-bad thus far?
"Doomsday predictions have occurred since time immemorial. Ergo, the Cuban Missile Crisis is nothing to worry about."
You actually have to look at the substance of the prediction. Sorry.
"I see your objection to ad hominem arguments, and raise you a further ad hominem argument"
I believe Altman is also a fan.
Altman says he thinks it is an "incredibly flawed movement"
https://x.com/sama/status/1593046526284410880
The dislike is mutual. Here's a long video takedown of Sam from a major EA org:
Just a few comments ago you argued that we don't know how to build superintelligence. Now you're saying we know how the (unevenly superintelligent) Fable system works.
It doesn't seem like you're being consistent here. I'm concerned there might be some motivated cognition going on.
"What is true is already so. Owning up to it doesn't make it worse. Not being open about it doesn't make it go away. And because it's true, it is what is there to be interacted with. Anything untrue isn't there to be lived. People can stand what is true, for they are already enduring it."
In this very thread I am being told that Fable is nothing but a bit of scale and refinement on well-known neural network techniques. And next I am told that we can't even imagine how to build superintelligence. Which is it folks?
Fable 5 doesn't represent anything new, other than scale and some refinement techniques, over the original LLMs.
Yet it is generating billions in revenue which Eliza did not.
Perhaps all we need is scale and some refinement techniques to eat a big fraction of the economy.
If unimpressive inputs lead to impressive outputs, that should make you more worried, not less.
Theorizing about nuclear winter is somewhat similar, in the sense of being inaccessible to experiment. Does that mean we should disregard the possibility of nuclear winter?
despite the incredible advances made in AI in the meantime
So the goalposts will be moved whenever necessary in that case?
No amount of incredible advances in AI will ever get skeptical HN commenters to take AI's implications seriously?
"The Gish gallops will continue until the nagging doubts have been silenced"
Imagine observing powered flight for the first time and saying: "This is religious fervor folks. I remember the mythological story of Icarus from when I was younger. Key word, mythological."
The existence of mythology describing Scenario X is not a valid argument against the plausibility of Scenario X.
If we can acknowledge the possibility of nuclear doomsday without running an RCT of sample size 100 Earths, 50 of which undergo nuclear Armageddon, to verify that nuclear holocaust indeed a real phenomenon... then we can do the same for AI. Understand the arguments being made instead of engaging in these guilt-by-association arguments.
Modern AI capabilities are already mind-boggling by the standards of 20 years ago. We should at least prepare for the possibility that trends continue on the current trajectory.
There's a difference between hamfisted/uneven attempts to enforce immigration law which is already on the books, and proposing a total freeze on legal immigration. Trump never did the latter. The US came closer to a total freeze in the 1920s.
Interestingly some Europeans also feel that Europe is becoming less free and are fleeing to the US:
https://www.foxnews.com/world/anti-greta-activist-flees-euro...