HN user

Bolwin

296 karma
Posts0
Comments193
View on HN
No posts found.

This is a lot of words to say "switch models instead of writing a plan file and starting a new session" which I do anyway.

That said, it takes me a while to reads plans and often I'll take a break, by when the cache has expired. At that point it may be better to start over

Grok 4.5 14 days ago

Yes Americans can do both, unless their boss dislikes it, but that applies the world over.

Tools come with a tool description in json schema format, but yes your point stands, it is not enough for opus 4.8 which I've also noticed having tool call issues.

Yeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), those stay forever.

In theory yes, but the average person can't really run the big open models.

This is already happening, try to find a provider that still hosts older, especially less popular or succeeded open models.

For me personally, I've been trying to access Kimi K2-0711. There seems to be only one provider left on openrouter (NovitaAI) and 3/4 requests error out

Hah, I noticed the same thing writing fiction with fable. Most models seem to go into a sort of "storytelling mode" where they forget their PhD level smarts. I had a character who is doing repair on a satellite. Most models would give you a half-baked explanation with some technical terms - half of them right half of them wrong.

Fable gave a description so deep that even I couldn't figure out what was going on and had to ask it to give me a simpler explanation.

It's pretty hard to measure because most context rot comes from related context and the model has to be able to figure which parts are truly relevant, which ones are relevant but stale, which ones to ignore etc.

Each relevant thing is basically a rule. Trying to so something with 500 rules is what's hard.

If you take a standard benchmark and just prepend a random book to it, it will not capture that

I don't use Claude Code. I use my own handwritten agent (formerly using Pi) and know every token that goes into it. There are zero memories to confuse it. The system prompt is 200 tokens and completely self consistent.

Plus I've found that the only time models go above 100k tokens anyway is when they've started looping at which point it's much better to go back anyway.

Anecdotally most models know their recall is terrible (or have been trained to act as such), that's why they constantly reread files before editing or while reasoning.

I see this said often and find it insane given how many times I find opus models making basic recall mistakes at <100k tokens.

Personally I consider < 60k to be the smart zone for opus. This is worse for opus 4.7 and 4.8 cause of the more granular tokenizer

FrontierCode 1 month ago

What is the "house" harness for minimax? They haven't released any

Am I Unc? 2 months ago

The social section is also just selecting for introverts it asocial people

But they expect a few wrong

MAI-Thinking-1 2 months ago

In my experience above 60k quality noticeably drops.

30k for open source models

Here's another: https://xcancel.com/FireworksAI_HQ/status/206010388602804673...

Fireworks is processing 30T tokens a day on open models, or about 210T a week. So about 40% of gemini? I'd say that's pretty good.

Two more points: 1. https://openrouter.ai/provider/fireworks ~5B tokens average daily on openrouter from fireworks, which is a ration of ~1:6000 2. https://openrouter.ai/rankings total tokens on openrouter, ~4T daily, and more than half seems to be open. Say 2T.

If other providers' ratio is anywhere close to fireworks, that's on the order of 10 quadrillion open tokens daily.

That said I'd guess the ratio is probably not nearly as high for most providers.

Unicode 18.0.0 Beta 2 months ago

I've never seen emoji used for subtext. Usually they just repeat or emphasize what's in the text