HN user

sailingparrot

6,031 karma

halcyon.cameo_3y@icloud.com

Posts9
Comments955
View on HN

Yes they distill, but if you think you can trivially get a frontier-level model by "just" distilling from Claude's public API. you fundamentally do not understand the amount of work that goes into a modern post-training stack.

Without even talking about the fact that any distillation that was done was on Opus, as the timeline of Mythos/Fable vs Kimi 3 release dates just do not match in any plausible way.

If you want to read an educated take from someone that has actually spent the last few years working on post training I recommend Nathan Lambert's: https://x.com/natolambert/status/2079616308203942332

When you have a couple hundred billion dollars on the line I have zero faith in the messenger

The issue with your reasoning, is that if/when an advanced AI goes rogue, it will necessarily come from a lab with a couple hundred billion dollars on the line.

So this is not a useful criteria to asses whether this is worth worrying about or not.

Hard disagree. You have to assume security guardrails can be by bypassed or will fail to detect an attacker. So if you are going to deploy this model in production non-airgapped, you better know how it will behave in this non-airgapped environment without guardrails.

And if you are too afraid to test it without guardrails, that probably means it shouldn’t be released.

Life needs energy to be moving around, without energy exchanges, by very definition, nothing interesting happens.

An inert element, for that reason is just not suitable for life. It's not a reasoning based on anthropocentricity it's just basic chemistry and mathematics. If things can't assemble together, and combine, and form more complex structures, you can't get life. If you could get life out of simple basic atoms, we would see life everywhere, and we would be creating it everyday in labs. We don't.

Doesnt mean life can't exist there by using other elements, but detecting helium is not increasing the likelihood of finding life there at the very least.

How much do you want incumbent multi-decade culprits to pay?

You are clearly not grasping the magnitude change in how many satellites we used to launch vs how many we are launching nowadays.

In 2026, we are putting 10x as many objects in space as we did just 8 years ago, with Starlink being the bulk of it: https://ourworldindata.org/grapher/yearly-number-of-objects-....

Starlink has 12.5k satellites in space and looking to ramp up massively, the biggest "multi-decade incumbent", oneweb, has 5% as many, about 600.

ChatGPT Work 13 days ago

I wonder why they haven’t simply continued to rebrand Codex as a general-purpose tool.

They have, if you try to download ChatGPT app, it actually downloads codex now, and the first screen is "Codex is now the ChatGPT App"

Why a randomized reservation order? [...] we wanted to create a system that would be less frustrating and more fair for everyone. A launch that starts at a specific day and time tends to reward bots, people with fast internet connections, talented gaming fingers for quick F5/refresh reactions, and those who can schedule their life around that moment. By accepting reservation signups over the course of a few days, without any incentive to be first, we're hoping to take away some of that friction.

This is nice.

Thinking is implemented as regular autoregressive generations by everyone, meaning its just regular tokens, but they appear between <thinking></thinking> special tokens which are then programmatically removed from what the user can actually see.

Idea somewhat similar to what you describe exist but they make steering/post-training/interpretation much harder.

MAI-Thinking-1 2 months ago

Yes and no. Yes from a user PoV, I don't really see a great reason to use this other than for enterprises that care about using a model not trained on copyrighted data (not sure what the market really is for this anymore, feels like this concern has been forgotten by most customers).

From a strategic PoV for MS, all the models you cited are distilling GPT/Claude/Gemini and wouldn't be anywhere as good as they are without this distillation, which in turn means you are dependent on OAI/Anthropic/G first shipping a good model to generate data for your training. This MAI model is trained from scratch with no synthetic data or distillation. So in term of benchmark its obviously much harder to get strong score and thus not a disaster if they can keep on improving.

« Rotary Actuators (The "Reflected Inertia" Trap) », « Quasi-Direct Drive (QDD) — The "Cheetah" Approach «

The pattern ‘something — The « metaphor » <qualifier> ‘ screams Gemini. Gemini seem completely unable to generate a section title that doesn’t follow this annoying pattern.