HN user

ishurand4

16 karma
Posts0
Comments15
View on HN
No posts found.
Claude Fable 5 1 month ago

The only one I see that thinks it is claude other than claude itself is the GLM series.

Claude Fable 5 1 month ago

Its a quadratic graph. It starts low but not that capable, gets better and more expensive, and then the time comes in which the capability needed is not the ones of the frontier models and then the price goes down on the companies who host the models that the capability is "good enough"

Claude Fable 5 1 month ago

Well, for me at least, I pay more for input (Up to 1M per prompt) than output (usually max 4k-8k)

Claude Opus 4.8 2 months ago

The numbers they show don't matter. "On multi-round coreference/context recall tests (often cited as MRCR or long-text retrieval benchmarks), Opus 4.7 reportedly dropped from roughly 78.3% down to 32.2% compared to Opus 4.6.", but what did anthropic do? They just stopped showing the benchmark altogether and then just show the cherry top ones that got improved on.

Claude Opus 4.8 2 months ago

And anyway, with quantum, there will be no need for frontier companies as you might be able to even run a 1T param model on a consumer quantum computer.

Claude Opus 4.8 2 months ago

They just showed the benchmarks it improved on but it regressed on so much more, such as the MCRR benchmark: "On multi-round coreference/context recall tests (often cited as MRCR or long-text retrieval benchmarks), Opus 4.7 reportedly dropped from roughly 78.3% down to 32.2% compared to Opus 4.6."