HN user

anon373839

4,133 karma
Posts10
Comments804
View on HN

All models are benchmaxxed, period. ”Jagged frontier” is the euphemism du jour, I believe?

Anthropic/OpenAI were touting PhD-level intelligence three years ago. And they’re still shipping models that aren’t smart enough to realize things such as the need to drive the car to the car wash (because they hadn’t yet hill-climbed that particular brain-teaser).

My impression is that, outside of the US, LLMs are not being marketed as an existential event by hucksters and death cult lunatics.

Kimi Work 2 days ago

It is, but I really think this project would be better as open source. They are already all-in on open for the most valuable part (the frontier model)! They might as well make the harness fully transparent so that questions about what's being accessed and transmitted can be avoided.

I personally wouldn't use any agentic harness from any company, American, Chinese, or otherwise, that is closed source. The Grok fiasco shows why.

Creating that kind of dataset takes an enormous amount of effort

Can you imagine the amount of effort it takes to write 15 trillion tokens worth of art, literature, source code, textbooks, scientific papers, news articles, etc.? No wonder Anthropic just scooped it up and took it for free!

Yet you don’t seem bothered by this. I wonder why.

Qwen 3.8 3 days ago

I think it’sa big, open question. There does seem to be a limit for knowledge compression at this size. But the behaviors that are learned in RL? It’s quite possible that they don’t actually require so many parameters. I was absolutely shocked when Qwen 3.5 was released and could perform reliably over 100-200k contexts with very limited hallucinations. It was a staggering jump in context-faithfulness from the preceding models of that size class.

May I ask you a personal question? What is motivating you to take up the frontier labs' cause in this way? Not a rhetorical question.

For my part, I'll happily disclose that I have an axe to grind. I think the major AI labs are an aggressive form of a cancer that's been ravaging our society. I want to see them fail, of course -- but more than that, I want to see the public develop an immune response to this.

I just can't wrap my head around why someone would expend so much effort speaking up on their behalf. They have, after all, highly compensated PR people doing that for them!

You skipped the part where Hyundai chose to sell the cars at loss in the hopes of eventually gaining a monopoly position.

Qwen 3.8 3 days ago

He recently did a walkback of that post. But ultimately, who cares? If the only way for AI to progress is in the hands of a few closed players, well, I don’t really think humanity needs that. Of course, it’s a preposterous claim in the first place. The ultimate reason deep learning and LLMs have made it as far as they have is the explosion of open research and research artifacts in the last decade.

Qwen 3.8 4 days ago

I’ve seen no evidence that he believes in anything. He comes off as just another slimy would-be monopolist to me.

nobody cares at all that it happened

Oh, no. I wouldn’t say that. If that happened, I definitely care: I’m positively delighted about it.

People infringe on Anthropics IP

Anthropic’s model outputs contain no IP. This is actually a simple legal proposition (rare in this field!) that derives from the fact that only specific classes of IP exist: copyrights, patents, trade secrets, and trademarks. Examining each, it is clear that API outputs do not qualify. Anthropic disclaims copyright in outputs; the outputs are not patented; the outputs are not secret (a prerequisite to having trade secrets); and trademarks are irrelevant in concept.

I strongly agree with the premise that distillation is not an “attack”.

But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena.

API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “cold start” problem faster. By far, what matters more is the quality and variety of RL environments the model learns from.

Out of curiosity, I just tried this exact question with Qwen 3.6. It proceeded to look through my git commit history and then gave me a summary of the changes committed on June 4...

Of course, yes, the models will provide propaganda-aligned responses to prompts that specifically mention certain political issues. I don't care for this behavior, but it's virtually never triggered in everyday use, and can be trained out if so desired.

so much of US economy is invested in AI.

This is a talking point that plays into the frontier labs' desire to be seen as "too big to fail". While yes, several hundred billion dollars have been invested in AI, (a) much of this is in the form of circular Monopoly-money deals, and (b) the US GDP is over $30 trillion annually. The real economy - the one that makes food, builds homes, provides medical care, etc. - is so much bigger and more important than the AI industry.

That's not to say that an (inevitable?) AI crash won't be the spark that ignites a big recession. We are well overdue for one.

The writing style, if not AI, is at least a bit tryhard.

Turning to the substance of the article: why do people feel the need to run this fast? I have certainly experimented with letting coding agents run amok. The first few times you try it, it feels like a superpower. Then you start examining the icky choices they made in a codebase that is now a dense forest. Then you have to expend a bunch of effort beating it back into submission. Or I guess you can YOLO and throw more AI at it, but then I agree with the person quoted saying "at that point, what am I still doing here?" This is not a satisfying or sustainable way to build, and there really is no reason other than hype and FOMO to do it.

Yeah. I may be naive, but I do trust the major cloud infra providers to offer real ZDR. Though admittedly, I haven't read their terms so it's possible that they also contain egregious loopholes.

Anthropic’s position being that it is entitled to train models on the creative works of anyone at any time, but its own slop generators’ outputs are sacred jewels that must be protected from being learned from.

I suppose this is like when Anthropic was using “prompt modification, steering vectors, or parameter-efficient fine-tuning” to poison the work of people working in the LLM field, including academic researchers.

I can report that it's working in oMLX. I've been experimenting with the ternary one; it is quite an impressive model! I've been grilling it on some deep learning/computer vision stuff and it's aced everything so far. Responses are thorough, accurate, sophisticated. General knowledge outside of CS doesn't seem as robust, which I expected. Honestly, I don't think the examples in the blog post do it justice.

That's obviously how the AI labs are trying to position themselves. But slop generators are not integral to anything. They most definitely should be left to fail, and if the market so dictates, the hundreds of billions invested should go to zero.