People don't realize the exponential diffictlty curve with juggling. The highest amount of balls ever juggled is 11.
HN user
helloplanets
The actual part on fine-tuning seems very short in the article. Did I miss a page where they have examples of fine-tuning it for different niche use cases?
Optimizing models to be fine-tuned is an amazing direction, but just makes me wonder how much better this actually is at being fine-tuned compared to other models. As none of the modern models are great at being fine-tuned afaik. Basically looking for some sort of benchmark showing that it's resistant to overfitting / catastrophic forgetting, etc.
Would be very interesting to see concrete demonstrations of different fine-tunes of the model. I'd imagine they've done hundreds of those internally.
Shouldn't the valuation be in Bs instead of Ms?
Yes, it is a 10x markup on the API prices. Depending on whether you factor in cooling costs, data center staff, etc. Or GPU costs and the electricity the GPUs are using only.
Either way, inference is very much where the money is made, training is where the money is lost.
True, it'd be a whole other situation if the tokens limits were cumulative. I guess it would all come down to whether their Claude Code subscription plans are turning in a profit or not.
At least for the segment of 20$ subscribers who actually use Claude Code it seems that it wasn't being profitable, as a couple months back they were testing out a pricing model where Claude Code would've not been included in the 20$ plan.
https://arstechnica.com/ai/2026/04/anthropic-tested-removing...
The subscription based plans are heavily subsidized, but the direct API inference pricing (which larger companies need to pay) is profitable.
Using a full Claude Max 20x plan to 100% of weekly usage would easily cost you 2k through the API. While the Claude Max 20x plan is 200 a month.
I just did this on one .claude directory and >20% of the answers there included some variation of "real", "actual", "exact", "honest", "genuine", "valid", "true". ~15% in that directory contain some variation of "real", "genuine" or "honest". This is excluding thinking tokens, sub-agent output, etc.
It's kind of offputting how much Anthropic models these days keep repeating "real", "genuine" and "honest". They've RL'd that way over the top.
Definitely not just a classifier layered on top, although there is one of those as well. Pretty sure it's different post-training / finetune run and the model weights are different between them. Which is bound to have effects on the model output even in general use, although I'd imagine they have pretty stringent evals to make sure it doesn't regress too much in things they don't want to restrict.
But this guy's been all over the news? He's not some made up fantasy person with absolutely no real world footprint. Even if his website is sloppy.
Not at all. This looks just like someone trying to make a quick buck, hyping their product up with bad benchmarks.
Both conditions used GitHub Copilot (Claude Sonnet 4.5 or Haiku 4.5, depending on study) running in VS Code within isolated Docker containers. The only difference was Mouse tool availability. (https://hic-ai.com/papers/mouse-paper-v13.pdf)
Haiku/Sonnet 4.5 on GitHub Copilot is not a valid comparison whatsoever.
You need to benchmark against Claude Code running Opus. I mean, being revolutionary is a big claim to fame.
These are my personal beliefs, not those of Nym.
Why are you posting this on your company's site, littered with ads for the company's product?
Post it on a personal blog, or just say that these indeed are the company's beliefs.
Doesn't make sense to fixate on LLMs and not the actual Transformer/attention foundation. The Transformer/attention architecture is the breakthrough, not LLMs. Especially the RLHF chat paradigm is 100% a byproduct. Which is easy to see when you look at how ChatGPT originally came about.
DeepMind has already has had real impact on science with the same foundational architecture as LLMs, for protein folding. They won a Nobel prize for it.
Tangential, but I'm pretty sad about EU having absolutely nothing in the actual SotA LLM market. Especially given the recent events of US completely restricting the actual SotA models.
Has this been just pure lack of funding and infra?
If that ain't getting steganographically tagged...
Dario's been openly talking how worried he is about China and labs getting synthetic training data off their models, for years. Most recently in relation to "Mythos level" capabilities.
Not really distillation, just synthetic training data.
The issue is that using Claude Code is an easy compromise for most to make, when you get to use the models 10x cheaper than through API pricing with a custom harness.
The cheap tokens are the product.
Slide number 55 is a beauty.
Which great writers are you thinking of here? True outsider art is very rare afaik.
I actually thought about that while writing the original comment as well. For Emma, Forever Ago is one of my all time favorite albums, good example of raw emotion with no need for any bells or whistles.
The big thing there is, that he already was a professional musician and completely inside a creative scene before leaving for the cabin. (DeYarmond Edison was the band he was in before Bon Iver.)
But yes, things were going way sideways for him, liver issues with mono, so he went to process whatever was going on and had been going on in complete isolation. (Although for the next album, he actually set up a whole "creative commune", a new band around Bon Iver instead of it being just himself, and so on. And I think you can hear the colors he wanted back in the music from it directly.)
A lot of examples of artists going into bouts of isolation, but almost always coming into it from an intense experience. So, the two don't have to be day to day intertwined, although for Techno specifically it's usually the case.
I don't believe this is how great music usually comes about, not even Techno. It's missing the other essential piece. Being influenced by and completely immersed in a niche of other brilliant people. (The most extreme example of this would be the 90's Detroit-Berlin connection.)
Paired with an obsessive work ethic in the studio.
If it's only obsession in the studio, things come out dry, uninspired. If there's no surge of energy running through your bones when making the music, why would anyone else feel anything? Mixing and the music sounding "professional" is completely secondary. Even detrimental a lot of the time, to be honest.
Applies to many other things than music as well. I don't any great technology comes out and about without that loop, either.
Deep Research has been using the Orchestrator -> Subagents -> Synthesizer loop since the beginning. It's just strange that they'd put a loop benchmark next to actual model benchmarks.
Maybe it's a tune of the base model that works especially well with the subagent loop?
OpenAI also announced two days ago that they're starting to make Cerebras style chips themselves [0], will be interesting to see how fast SotA model inference will be by the end of the year.
[0] https://openai.com/index/openai-broadcom-jalapeno-inference-...
Less so in EU than in US.
I'm not saying your actual point couldn't be valid or fully defensible, just to be clear.
My view is that there are people capable of vetting LLM generated code, and people who are not capable of it, based on their previous track record of vetting non-LLM generated code and the quality of their own non-LLM generated code.
For example: I would trust the capability of John Carmack to vet an LLM generated bug fix, to his own game engine. Even if it was LLM generated by him, and vetted by him.
Slop means anything produced en masse with complete disregard for truth, accuracy, or usefulness.
This doesn't match at all with what the author described in the article.
Anyone trying to say "but my slop isn't slop, I vetted it" clearly is not in possession of the necessary critical thinking skills to differentiate between slop and non-slop.
This is called a Kafkatrap. It works in any direction, in any situation, making the disagreement moot. Also not considered good faith rhetoric.
That is not what slop means, though. You're redefining the meaning of the word to suit your view. Why do that? You can just say that LLM generated content is not up to par, or acceptable, ever.
Maxed out 2019 Mac Pro was $50k+. The wheels on that thing were $400. This is a bargain compared to that.
Anthropic already explicitly communicated that they'll store and check all the data from Bedrock or any platform, even if you've selected zero data retention, if using Mythos class models. To use these models on any platform, you'll have to accept these terms regardless of the region.
Limited data retention and review as part of our safety work. Prompts submitted to, and outputs generated by, Mythos-class models are retained for 30 days for trust and safety purposes, on every platform where these models are offered.
Change applies to organizations that have set up workspaces with zero data retention (ZDR) in Claude Console, use Claude Code with ZDR in Claude Enterprise, or access Claude through AWS Bedrock, Google Cloud Agent Platform, or Microsoft Foundry with ZDR.
https://support.claude.com/en/articles/15425996-data-retenti...