theoretically possible
I mean I guess, but not in a performant way if there are ever any hardware failures. And with 100K GPUs there are multiple hardware failures per day.
HN user
Formerly founder of an ML company in Adelaide and Sydney. Exited 2021.
Did some decentralized databases work and private ML work.
Formally CTO at a startup working on AI for strategy/policy work.
Always happy for emails or tweets at me.
http://twitter.com/nlothian
firstname.lastname at gmail.com
theoretically possible
I mean I guess, but not in a performant way if there are ever any hardware failures. And with 100K GPUs there are multiple hardware failures per day.
Go turn on min_p once it's available post July 27th and most of the problems you describe will go away.
This seems both arrogantly dismissive ("you are holding it wrong") and incorrect.
Either the OP is using Kimi K3 on Moonshot where is is presumable set correctly (K3 isn't available elsewhere yet), or they are using Kimi K2.x and there has been plenty of time to experiment with this.
set up the environment sloppily
The model used a zero-day exploit to escape, and then multiple chained privilege escalations to escape.
That indicates the environment both was hardened against all known attacks and had defenses in depth.
wildly negligent to not be running it in a physically-airgapped environment
Why should it be physically airgapped? Clients won't be doing that.
I don't understand this sentiment at all.
Is it a claim that "breaking into Hugging Face's production infrastructure" didn't happen? That it's not actually all that severe? That it was done by hand by OpenAI employees and they fooled Hugging Face?
That the blog post exaggerates something, somehow?
What exactly do you mean?
At the moment it just reads like a thoughtless dismissal.
Deterministic seeds barely work on a single machine, small scale training run.
They just don't work at all on a many month long, 100K+ GPU cluster training run.
Great, but that seems a different concern to the auditability of a model.
You can take the code for Kimi K3 now, take the training framework from Prime and the data from Olmo, spend some money on RL environments and some more money (!) on GPU training and end up with a system of similar capabilities.
But that's completely different to being able to audit Kimi K3. Even if you had the exact code, data and training environments it is impossible to verify that the model you have came from that.
I know someone who runs AI training.
They will have people who don't understand the distinction between visiting Claude.ai and downloading Claude Cowork.
They type the words "setup MCP" into Claude.ai and expect it to automate Excel on their machine.
There's a pretty big gap between the things we talk about here, and where the world is at.
Getting a SOTA model in a custom harness requires API pricing or risking an account ban, correct?
No, only Anthropic has that policy (and I think even that is relaxed for an unknown period if you use the Claude Agent SDK: https://support.claude.com/en/articles/15036540-use-the-clau...).
OpenAI, Kimi, Qwen, GLM and Deepseek all allow it.
I'm not sure about Gemini.
because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?
Is this an assertion that is backed by evidence?
From the Elon/OpenAI trial:
On the stand in a California federal court on Thursday, Elon Musk was asked if xAI has used distillation techniques on OpenAI models to train Grok, and he asserted it was a general practice among AI companies. Asked if that meant “yes,” he said, “Partly.”
https://techcrunch.com/2026/04/30/elon-musk-testifies-that-x...
Have you ever worked with a non-programmer and helped them setup their AI workflows?
You install MCP connectors, specific skills, work around model/harness quirks, set security boundaries etc.
It's a lot of work, and most people will never want to change it once they have it working.
I'm all for open models, but people seem to misunderstand what they are. They aren't the same thing as open source code!
open weights, open code and open data
Even if you have all these things you still can't replicate a model because of randomness.
You can backdoor a model with less than 1000 examples and it is impossible to detect.
China's strategy of spending billions on training these models and open sourcing these models away is strategic - they want to kill the US LLM industry at any cost.
Why is it when Anthropic and OpenAI spend billions trying to beat each other it is competition, but when the Chinese companies do it then it is trying to kill the US LLM industry at any cost.
The US federal government spends billions in subsidies via the US Chip Act, and bans chip sales to China to support US companies.
But the implication is that somehow Chinese competition is illegitimate because "strategic".
That's just not true. You can absolutely have terms of service that are illegal, and the government can enforce them.
Sure, but her writing wasn't in any way political.
having PhDs doing night shift lab tech work for pennies
I don't know why people keep bringing this up as though it is surprising.
In almost any field other than AI PhDs are underpaid on average.
There are many, many bio PhDs working as lab technicians.
No reason to think it will. Paul Allen died after donating the money to create AllenAI and I don't think there are any links.
Hopefully it somehow works out though!
Meta isn't investing in frontier big models anymore
Yes they are. Meta Muse is their attempt.
It's below frontier performance at the moment but they are spending on getting there.
Llama 4 was a bad architecture.
Meta Spark is moderately promising but of course closed source.
AllenAI is great, but they don't have the budget or remit to build large models.
If the path that was taken to commit this is full of "oops" and "fix" messages
great way to encourage people to rebase then!
Undecidable isn't uncomputable.
"Computable" can mean probabilistic, and classical computers can function over probability distributions just fine.
I wouldn't be too fixated on the specific numbers in that post.
Anthropic was extremely capacity constrained at that point. They still are but not to that extent.
I'd note that OpenAI offers 24 hour caching. I'd be surprised if Anthropic hasn't optimised their caching for Claude code too.
SemiAnalysis recently posted that their actual Opus usage works out at $0.99 because of caching.
The principles remain though.
It's not one prompt, but here is a parametric rod connector:
Use SCAD and design a connector for square rods.
The rods are 18.2 mm square. I want to connect two end-to-end.
.. make if the bolt holes are created optional for each side - I want to set them separately. Make them M3.5 countersunk
.. it's the +X or -X sides I want to turn the screw holes on or off.
.. on each of the 4 sides of the connector add additional connectors at 90 degrees. Make each optional
etc etcI think that applies to military involvement abroad generally.
If you are dropping bombs on someone I'm unconvinced the use of AI will make them like you more or less.
I've been doing a lot of 3D design in Codex GPT 5.5 (I found Opus 4.7 wasn't as good - haven't experimented much with 4.8 or Fable).
OpenSCAD is a parametric CAD programming language, and the models know it well.
The biggest challenge is communicating words like "inside" and "above" to the model - inevitably it's idea of which direction is which is often different.
I can't say I've done anything very hard, but for things like ESP32 cases, or parametric rod connectors it is great.
You can do things like "add snap connectors" and it'll do a great job.
It's been very successful at frontier math tasks - a bunch of the Erdos questions have been solved by it - more than any other model.
Their methodology isn't published.
Its widely accepted[1] that it runs the same query through the model in parallel and then has a model that either selects the best answer or synthesizes an answer from the multiple ones generated.
I believe most people think it runs 6 sub-models, but I think that is based on the pricing.
It's a pity that OpenAI doesn't publish details like this.
The source is the GPT 5.5 System Card:
We generally treat GPT-5.5’s safety results as strong proxies for GPT-5.5 Pro, which is the same underlying model using a setting that makes use of parallel test time compute. As noted below, we separately evaluate GPT-5.5 Pro in certain cases because we judge that the setting could materially impact the relevant risks or appropriate safeguards posture.
https://deploymentsafety.openai.com/gpt-5-5/model-data-and-t...
There have been multiple podcasts with people from OpenAI which have confirmed this.
I like this format:
"I love Lean because <abc>. I found it failed in <xyz> case because <123>. I created a thing <blah> which handles that like this: <ahhh>.
I'd love feedback! It's open source here: "