I think they're distilling "reasoning", not merely code. There's plenty of human generated code. What they're trying to extract is the process by which really large models "think" their way through complicated problems. That's what all the "traces" datasets on HuggingFace are about.
HN user
SwellJoe
YC W07
I work on Virtualmin. You can find it at http://www.virtualmin.com
I blog at https://swelljoe.com
I toot at https://mas.to/@swelljoe
"It is already understood that you own the output."
I don't think it's settled that anybody owns the output. There seems to be some question whether LLM output can be copyrighted (and there should be).
I'd rather it weren't possible, actually. I think it's better for humanity if we acknowledge that what was legitimately ingested into these models is our collective commons (and what was illegitimately ingested into these models also shouldn't exclusively profit the people who illegitimately did so). I don't know how that squares with the AI industry recovering its trillion dollars in investment, but I reckon they should have thought of that before.
Every bike is a balance bike. That's how bikes work.
This is exactly the kind of model that's been needed in the middle. Realistically self-hosted, Good Enough intelligence, MoE so it's fast on limited bandwidth systems like Strix Halo and DGX Spark.
For a while there's been nothing to run on my Strix Halo that's notably better than what I can run on my dual 32GB GPU desktop (Gemma 4 or Qwen 3.6 dense models), but this seems likely to be the step up in size that actually works better than those.
It's also got vision and audio. So, the better comparison is any of the other large Chinese open models that are better and cheaper than Gemini Flash.
3.5 Flash was always too expensive for a "flash" model. They marketed it as "near frontier" level, but there are several order-of-magnitude cheaper open models that compete with it.
I'm not willing to believe in phone-based frontier models anytime soon. Though, Gemma 4 12B is a beast that runs comfortably on the current top of the line phones (or would run fine if allowed to run, I think there's some kind of 6GB limit on iOS, and 12B is ~7GB). I'll believe in three years we'll be able to run ~30B models on the best phones. That's 16GB in a 4-bit quantization, and I believe ~30B models will be competitive with 120B models of today, based on the curve we've been on. Qwen 27B and Gemma 4 31B are competitive with much larger models of a couple years ago.
"I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now."
I don't think it'll take 10-15 years. Gemma 4 31B in the 4-bit QAT is competitive with the frontier of less than three years ago and runs on any high-end 32GB gaming PC GPU or a large-ish Mac.
The question is whether the frontier will continue to get better at a rate that allows it to stay ahead of the two curves of availability of consumer hardware big enough to run somewhat larger models and the capability of small models to compete with large ones. When the bottom falls out and GPUs/RAM becomes affordable again, the size of what normal people have on their desk will trend quite a bit larger than today.
I think there's a future not too far from now, where a 120B model with really good reasoning and a large context, but limited knowledge (necessitated by being small, you can't fit the world's knowledge in 100 gigabytes), can substitute for a frontier model on almost any task, just by giving it access to web search and documentation for the thing you're trying to do. A 256GB unified memory machine with sufficient memory bandwidth would comfortably run that 120B model.
Even very large models failed tests like this up until a year or two ago...this is a very, very, small model.
I've been seeing more "grounding" messages in the thinking log of Claude models. I guess that's "do a web search, read the docs", probably designed to mitigate hallucinations.
I am surprised by how much the really giant models hallucinate, though. My vague feeling was that little models hallucinate a lot because they just don't know anything (the world's knowledge simply does not fit in a few GB) and don't know how to say, "I don't know". But, the big models kinda do know everything, and yet, here we are, they're still making shit up all the time.
In my benchmarks of security auditing abilities of models, DeepSeek was roughly an order of magnitude cheaper than either GPT 5.5 or Opus 4.8, less than ten cents per task vs. roughly a buck each for the best American models at the time.
GPT 5.5 Pro was ~230x at almost $23 per task.
DeepSeek is my go-to when I need an API, and local Gemma 4 won't do because it's either too slow or not capable enough. DeepSeek isn't at the frontier but it's good enough for a lot of things, very cheap, and quite fast. Flash is even faster and cheaper, and still better than anything I can host locally.
You still don't know what's going on in there.
That's interesting. The "best" current models, Fable and especially GPT 5.6, are also lying liars that lie all the time. Seems like we're going the wrong way on hallucination.
Nobody does caching as well as DeepSeek, so I guess it's a big enough difference in the implementation to make it difficult.
If you use Reasonix with DeepSeek it gets silly, as it is append-only to work with how caching works. It gets something like 97-98% cached tokens in a long session. It makes an already cheap model even cheaper.
I find the 4-bit QAT with MTP to be entirely usable speed on both my boxes (Strix Halo and a desktop with two V620 GPUs, which are slightly faster than the Strix Halo).
I have seen occasional weird behavior that I guess could be attributed to hallucinations, but for security auditing, DeepSeek v4 Pro is among the best models I've tested, competitive with Opus 4.8 and GPT 5.5 (MiMo and GLM also did well, Qwen 3.7 Max was below all of those, though only barely), and at an order of magnitude lower cost per task.
Qwen is the most censored of the Chinese models in my testing, which makes me wonder in what other ways it is compromised. Open weights doesn't really reveal what's in there. And, in my tests, existing Qwen models are not at the pareto frontier of any metric; DeepSeek V4 Pro is better, faster, and much cheaper than Qwen 3.7 Max. (DeepSeek is also among the least censored of the Chinese models.)
I guess we'll see if the "second only to Fable" hype pans out. In my limited experience with Kimi K3 (I signed up for a month of the $19 plan) it's slower and chews a lot more, so ends up being pretty expensive; one little feature burned through almost the entirety of my five hour limit. The $20 GPT plan is a lot more useful and includes 5.6 Sol, which is fast and token-efficient enough to be quite usable even with the small plan.
Kimi K3 weights aren't available yet. And, the claim is about Inkling. I'm not saying it's wrong, I just don't understand why Inkling would be the trigger, rather than a prior model that's better, at least for code. Seems like a general trend of continuing scarcity and people finding new ways to use models, so usage just keeps climbing without sufficient supply. Everybody wants more GPUs and there aren't enough being made.
I'm not sure I buy the direct causation? Why didn't GLM 5.2 have that effect? It's a better model, by some measures, and a little easier to run as it's quite a bit smaller.
That sounds complicated. I'll just use my month of Kimi and then cancel. I have too many AI subscriptions to use them all, anyway. I subscribed mostly to test it. I mean, if it turned out to be competitive, I would keep it, but if it doesn't turn out to really excel and anything and also take longer than Claude or OpenAI models, I'll stick with them.
It let's me choose different thinking levels in Kimi Code. Not sure if it actually works, yet, but it says "Thinking set to high." when I change it from max.
Yep, with Reasonix, DeepSeek is free real estate. Seems to just go and go for pennies.
And, DeepSeek is what I use for any task that works best with an API. It's cheap enough to where I don't think about cost, made even cheaper by DeepSeek having the most effective and cheap caching in the industry, and it's good enough to where I rarely have to follow up with a more expensive model or manually fix things. It's been alleged they're releasing an update to DeepSeek V4 Pro soon that improves it, which likely makes it a good fit for even more kinds of problems. It remains my favorite of the Chinese models, it's so cheap and cheerful. And, is also less aggressively censored than some of them.
Yeah, I'm finding I end up switching to Codex and GPT 5.6 a lot lately because I've either run out of Fable usage or Fable refused to do the task. Most recently it refused to work on a WiFi configuration UI for a robot. No idea why it thought that was related to security, biology, or some other sensitive topic. They've hobbled it with guardrails that are overzealous and now there's a big opening in the market. Fable may be the best, but if it won't do the job half the time, it stops being my go to model as I don't want to waste time only to find it refuses halfway through.
Not sure how the economics work for the Chinese models, but DeepSeek did the same task for a dime.
I tried Kimi K3 on a task I've done with every other model I use regularly (https://swelljoe.com/post/i-let-every-agent-implement-its-ow...) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan.
I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed almost none of the 5 hour limit.
Subscription usage limits are hard to measure as none of the providers tell you directly what it means in terms of tokens or anything else you can easily compare, but when I sat down to add Kimi Code to flar, it was because I wanted to try it on some real work and then couldn't do any, because usage was nearly gone after the trivial task...no other ~$20 subscription I have has felt that tight before.
So, it was really slow to complete the task and seemingly much more expensive than every other model I'd tried. Maybe bad luck. Maybe it'll do better on other tasks. I wouldn't know as I was out of usage when I had time to try.
It did find a bug that Gemini 3.5 Flash introduced unprompted, though, so it has that going for it.
They're being generous.
Investors are clearly not familiar with anything Musk does, because it's all wildly overvalued.
A dip because of a scrubbed launch is a blip on the radar compared to the catastrophically bad ideas Musk is promising to implement. Data centers in space? That's just lunacy. A ridiculous idea from a ridiculous man.
This is why I'm not wasting my time on these things (either in participating or in paying attention to the results). The noise is just too high. I'd be furious if I'd spent a bunch of time on doing actual work for something like this only to have slop win. Hopefully, the other contestants worked on stuff that is transferable to other purposes.
I buy DeepSeek and MiMo directly from them, on the assumption they're better at it, have a vested interest in me having a good experience and they'll get caching and other stuff right. They're also cheap enough that it's unlikely you'll find a much better deal on those models.
I'm likely to also add a small Kimi Code subscription, as the model looks very promising and I don't see any reason to support proprietary US models from companies I don't really trust overopen Chinese models. I've opted not to get a z.AI subscription, though, as the price/performance ends up not being great, because GLM chews a lot more on the problem so its actual price per task is roughly the same as the big guys. The same can't be said of DeepSeek. It's notably cheaper per task, especially when comparing API rates to Anthropic or OpenAI.
I think that's probably a good thing. Sycophancy seems to be correlated with AI psychosis. GPT 4o was creepy sycophantic and has a body count. It'll be good for chatbots to be more interested in facts than in agreeing. (Then again, I found Qwen 3.6 to be strident in its lies about Uyghurs in China, among other "sensitive" topics, parroting the party line and getting almost hostile when told to search the web for current information.)