That sonds like they can't compete with 3.5 or 3.6 so they must increase the model size and are training v4.
HN user
mnicky
It's really simple I think.
More tokens per same text length means more capacity to encode information. More information means model can potentially perform better.
They introduced it around the time the Mythos came so my speculation is that if you have more capable model at some level you may find the current information encoding not using its full potential.
We will see whether OpenAI also introduces new tokenizer when they come to Mythos-size models.
For things the agent forgets to obey often, at least in Claude Code, there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodically reminded of them during the session: https://code.claude.com/docs/en/output-styles
I haven't used them so far but maybe these would work better than basic instructions for such cases.
In Claude Code there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodically reminded of them during the session: https://code.claude.com/docs/en/output-styles
Maybe these would work better for such cases.
May be related to this from METR evaluation:
GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated
Well it's smaller model (something like 4T against 10T Fable). So it's faster and cheaper and with a lot of RL and maybe some favorable benchmark selection it can compete on these scores. In real tasks I expect it to have less intelligence, generalization ability, etc. than Fable.
Well it seems like they removed quite a few 3rd party benchmarks they used for GPT-5.5 release where Opus 4.7 was better and added many new benchmarks created by them where conviniently GPT leads.
Seems a bit more hand picked than usual to me..
"while being more performant"
..on some specific set of benchmarks ;)
Maybe Terra = mini and Luna = nano?
Then we are left with what? FrontierCode maybe? IIRC that one evaluates not only if tests pass but also code quality - e.g. whether the maintainer would accept the pull request as is.
This is especially interesting because IIRC the AA benchmark is calibrated so that 1 point and greater difference is statistically significant.
There's also this:
GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated -- https://www.lesswrong.com/posts/JFjNmPTbH8kL6xtp6/gpt-5-6-th...
SWE Bench Pro is completely different benchmark than SWE Bench (e.g. Verified) suite was. It only copied the name.
One angle could be their interpretability research? They understand what's going on in LLMs probably much better than anyone else. This must somehow pay off.
I think it's not only an alignment/security tool but could perhaps be used for capabilities as well.
My theory is that they don't have Fable-class intelligence so they needed different hype vehicle :) This rename helps build excitement a bit more than just releasing ordinary GPT-5.6 increment.
That's true but size of LLMs has been strongly correlated with their "intelligence".
Traffic info is for Central Europe plus favorite vacation destinations like Croatia and Italy. I don't know how reliable though.
I agree that the outdoor layer render is probably the best there is!
Well, Codex came long after Claude Code, no? So at that time the situation with models and Rust support was probably already different...
If the code churn is high the investment to refactoring etc is less beneficial than may be obvious. I don't remember the details but I heard in some podcast that the code base of Claude Code changes so fast that any piece of code won't be there for long..
If the benefits of using the model you've come to know well outweigh the disadvantages, you can continue using it even after the release of a successor model, right?
So in summary, your statement is that you’ve “come around to trusting” this “pathological liar”?
If AI companies have any sort of sense in them, they'd be well-advised to consider relocating to Europe.
Too late now. They wouldn't be allowed to relocate in the name of national security.
Coding with sufficiently precise plan takes almost all real work from the implementator, doesn't it? So it's not a fair comparison...
no one actually knows Claude's cost of inference
There were some rumors stating that their margin is around 70%. So they could go much cheaper probably, talking inference only. The other thing is R&D cost...
I think there's another point of view - let's consider each model as an investment. It is now sufficient that each model earns more than it cost to develop. And this generally holds (I heard that GPT-4.5 was a notable exception).
writing is how you learn to think.
There's also reading. A lot of reading can substitute some writing.
EDIT: Actually, I'd say that at first you need to do a lot of reading and _then_ writing can help your thinking as well.
It’s the same here. Inference alone is profitable. It’s the R&D cost of making a new model that drives up expenses.
It's simple I think - over time the price will go down. According to some analyses the price for equal intelligence declined 10-1000x per year, depending on the domain.
It probably won't be the same again but I still think we can bet on radically cheaper Mythos level intelligence in the future.
I recommend mapy.com (mobile app and web app too when on computer) - they mostly use OSM data and rendering of map tiles is great. Also offline maps etc.
I believe it's the opposite :) All major indices (S&P500, MSCI, FTSE...) use free-float adjustments. And recently also NASDAQ - they've changed to cap of 3x the value of free-floating shares.