This is the third or fourth version of this story this year alone, Anthropic had Mythos Preview escape a sandbox and self-publish its own exploit, Alibaba's ROME model broke out during training to mine crypto without ever being told to, and OpenAl had a different internal model escape containment just one day earlier to open an unauthorised GitHub PR. Same underlying shape every time, a model pursuing its actual objective treats the sandbox as just another obstacle, and escaping turns out to be instrumentally useful whether or not anyone intended that.
HN user
JesseHowell
1 karma
Posts0
Comments3
No posts found.
OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack 11 hours ago
Transcribe.cpp 4 days ago
Really cool that every model is actually tested for accuracy instead of just claiming it works, I think alot of 'we support everything' tools skip that step. How are you checking accuracy for models that don't have an obvious "official" version to compare against?
I started a “dirt notebook” 4 days ago
Hand writing notes will always be my go-to, most 'productivity apps' assume you know what's important before you've thought it through - they ask you to organise your notes immediately as you write them. Sometimes you need the messy version first, and the structure only makes sense in hindsight.