HN user

mnicky

111 karma
Posts6
Comments89
View on HN

It's really simple I think.

More tokens per same text length means more capacity to encode information. More information means model can potentially perform better.

They introduced it around the time the Mythos came so my speculation is that if you have more capable model at some level you may find the current information encoding not using its full potential.

We will see whether OpenAI also introduces new tokenizer when they come to Mythos-size models.

GPT-5.6 13 days ago

May be related to this from METR evaluation:

GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated

GPT-5.6 13 days ago

Well it's smaller model (something like 4T against 10T Fable). So it's faster and cheaper and with a lot of RL and maybe some favorable benchmark selection it can compete on these scores. In real tasks I expect it to have less intelligence, generalization ability, etc. than Fable.

GPT-5.6 13 days ago

Well it seems like they removed quite a few 3rd party benchmarks they used for GPT-5.5 release where Opus 4.7 was better and added many new benchmarks created by them where conviniently GPT leads.

Seems a bit more hand picked than usual to me..

GPT-5.6 13 days ago

"while being more performant"

..on some specific set of benchmarks ;)

GPT-5.6 13 days ago

Then we are left with what? FrontierCode maybe? IIRC that one evaluates not only if tests pass but also code quality - e.g. whether the maintainer would accept the pull request as is.

GPT-5.6 13 days ago

This is especially interesting because IIRC the AA benchmark is calibrated so that 1 point and greater difference is statistically significant.

GPT-5.6 13 days ago

SWE Bench Pro is completely different benchmark than SWE Bench (e.g. Verified) suite was. It only copied the name.

Grok 4.5 14 days ago

One angle could be their interpretability research? They understand what's going on in LLMs probably much better than anyone else. This must somehow pay off.

I think it's not only an alignment/security tool but could perhaps be used for capabilities as well.

Traffic info is for Central Europe plus favorite vacation destinations like Croatia and Italy. I don't know how reliable though.

I agree that the outdoor layer render is probably the best there is!

If the code churn is high the investment to refactoring etc is less beneficial than may be obvious. I don't remember the details but I heard in some podcast that the code base of Claude Code changes so fast that any piece of code won't be there for long..

writing is how you learn to think.

There's also reading. A lot of reading can substitute some writing.

EDIT: Actually, I'd say that at first you need to do a lot of reading and _then_ writing can help your thinking as well.

It's simple I think - over time the price will go down. According to some analyses the price for equal intelligence declined 10-1000x per year, depending on the domain.

It probably won't be the same again but I still think we can bet on radically cheaper Mythos level intelligence in the future.

Siri AI 1 month ago

I recommend mapy.com (mobile app and web app too when on computer) - they mostly use OSM data and rendering of map tiles is great. Also offline maps etc.