HN user

sigmar

5,124 karma
Posts2
Comments1,042
View on HN

We are currently working closely with our inference partners and open-source maintainers to align the technical details and ensure the model can be reliably deployed across the ecosystem. The full model weights will be released by July 27, 2026. Further details regarding the architecture, training, and evaluation will be released with the Kimi K3 technical report.

(translated by chrome)

11 days is a long time. It does not take that long to implement inference at providers. In my opinion, seems like they're being pre-emptively cautious about government intervention/review

Most of what an LLM does "could have" been done by a human if you throw enough human hours at it. But the reality in this circumstance is that a new tool helped find this leak. Saying this could have happened in a "non LLM world" is analogous to "someone else could have discovered special relativity, let's not mention Einstein"

Evaluators validate each tag individually — for example, protein, preparation, or health, individually rather than judging the item as a whole.

Am I reading this right that the jury is multiple LLMs each iterating through each tag and voting on each? Why wouldn't you tune one LLM to be really competent at a single tag? Like a single "spicy evaluator LLM" or "protein evaluator LLM"?

To me, this addendum makes it worse. Making small edits to a post like this makes it seem like you're doubling down on the original resentful points, especially with all the new justifications like "a trillion dollar company fired the first shot." Should have just deleted the blog post in its entirety imho.

outstanding relationships with essentially everyone who openly uses Zig and talks about it publicly

Essentially everyone? more blogposts incoming?

One, this constant bullshit about some window closing, or the perpetual underclass, or falling hopelessly behind. This is negative valence hype, not only is it not true, it’s mostly designed to make you feel bad about yourself and move to shitty San Francisco where everything really does suck like how these people claim.

It's possible to use LLMs without logging onto twitter to be exposed to the people spouting off about a "perpetual underclass." I love the internet, but it really feels like (now more than ever) you have to be intentional about what sites you visit.

They don't make any specific claims about what conditions it will diagnose. At 16:30, he says they are only initially doing "body composition" because anything more would add 9+ months to the deployment timeline. I assume they mean they process the images to return an estimate of body fat/muscle mass. Which isn't difficult and it seems likely they could get the error bars pretty low just off estimating subq fat alone. He doesn't say the specific classification they received from the FDA, just that it is a class 2 medical device.

I visited CERN last July. Was lucky enough to get into a group tour. The tour guide was a postdoc researcher who said the only times that public tours are allowed to take an elevator down is during long shutdowns. So while they do this work on LHC might be the best time to swing by for a tour (I might even try to return).

Even without the descent, my tour was great with showing the 70 year timeline, historical early particle accelerator equipment, and a cool view of the ATLAS control room. The facility is awe inspiring and a testament to Europe's willingness to make long-term commitments to furthering science research for the public good.

publish these incredible papers explaining how they achieved their gains - something the American labs no longer do unfortunately.

Google is still releasing a lot of llm architecture research. They introduced speculative decoding of LLMs in 2022[1], then released the code to perform sceculative decoding for their Gemma 4 model this year[2]

[1] https://arxiv.org/abs/2211.17192

[2] https://github.com/google-gemma/cookbook/blob/main/docs/mtp/...

Qualified it with "100%" because claude4 models show the first few lines of the chain of thought:

On Claude 4 models, the first few lines of thinking output are more verbose, providing detailed reasoning that's particularly helpful for prompt engineering purposes. Claude Mythos Preview summarizes from the first token, so its thinking blocks do not show this verbose preamble. https://platform.claude.com/docs/en/build-with-claude/extend...

Some administration officials have said that a resolution should include an acknowledgment on Anthropic’s part that its rollout of Fable and communication with the White House could have been improved, people familiar with the talks said.

followed initial frustration Friday among some administration officials when they couldn’t immediately get Amodei on the phone, the people said.

That he didn't drop everything to talk to them seems like the major crux? But Dario doesn't even do the day-to-day operations Daniela does. Feel like Anthropic should just hire Dean Ball to be their liason or something

I should have contextualized the quote- "chat is dead" is from an openai employee which was describing how they're shifting focus to more agentic consumer products, and putting less focus on the back-and-forth chatbot interface.

I like that "chat is dead" framing I heard recently because too many people are having interpersonal relations with these LLMs and want to tune their "emotions"/tone. Humanity would be in a better place if we thought of the LLMs as tools and not friends. (even though they are very good at beating a turing test)

GLM 5.2 Is Out 1 month ago

Did you read the blog post where they explained why there was a temporary block on all biology-related questions?

Agree with this. Strange to me to frame the "training recall" as cheating (33 of the 38 cheating instances). Most people think of "cheating" as breaking rules. How is the LLM model supposed to not use what was put into the weights?

As long as Codex remains so affordable and useful they do not have to slash prices, just keep Codex usable.

I imagine they track usage and can see whether their habitual users are switching to something else and aren't going to slash prices 'for the hell of it'.

just look at public stats on openrouter (obviously not indicative of first party app usage or direct api usage, but there's a huge difference between these graphs): https://openrouter.ai/openai https://openrouter.ai/anthropic

It's temporary. From the fable blogpost:

To release the model both safely and quickly, we’ve tuned these safeguards conservatively—they’ll sometimes catch harmless requests, though they trigger, on average, in less than 5% of sessions. With more capable models arriving in the coming months, we’re working to improve our safeguards and reduce false positives as quickly as we can.

Claude Fable 5 1 month ago

The system card is 319 pages, at what point do we call it a "book" instead of a "card"?

There's a quote from a METR report on page 52:

We ran [Mythos 5] on 38 of our hardest software tasks, including tasks centered around R&D. [Mythos5] generally outperformed an early checkpoint of Claude Mythos Preview in these, including by succeeding on some tasks that had not been solved by any public model we have previously evaluated. However, we still observed the model occasionally failing to correctly interpret nuanced instructions in difficult tasks... Based on the available evidence, we believe [Mythos 5] is likely unable to fully and reliably automate R&D for frontier projects spanning multiple weeks. We believe that a better, more confident assessment would require more time, evaluations, and information from the model developer.