HN user

Topfi

2,893 karma

Most of my comments are, unfortunately, subjectively and in my personal opinion somewhat overlong due to my, admittedly not necessarily problematic tendency to hedge my own statements which, while in some cases helpful to more clearly delineate when I am referencing empirical data in contrast to when I am making a statement based on my personal and somewhat limited, due to my life experience at this age and the fact that I have grown up in the EU, experiences, I have grown to find that I have on occasion taken up the habit of overindulging on said clarifications and thus have come to the conclusion that this may be something I should address, as such extensive clarifications can make it draining to read my writing and can hide my actual intended point inside an avalanche of unnecessary and annoying to read writing, which is why in more recent comments of mine, which I have written and am intending to write in the future should time permit and topics I am interested in arise, you may find a more direct, albeit still where necessary properly clarified, style of writing comments...

Posts169
Comments471
View on HN
aistudio.google.com 1d ago

Gemini 3.6 Flash released on AIStudio

Topfi
7pts0
www.youtube.com 3d ago

The Newest Vegas Loop Stations Will Blow Your Mind [video]

Topfi
3pts0
basaltlabs.org 4d ago

Basaltlabs Monolith-1.0 – #1 on Last Exam, AIME, GPQA Diamond, MMLU-Pro

Topfi
1pts2
www.eu-searchperspective.com 4d ago

European Search Perspective

Topfi
1pts0
www.youtube.com 7d ago

Microsoft Project Aion (Copilot OS Incubation Effort) [video]

Topfi
3pts0
thinkingmachines.ai 7d ago

Inkling Model Card

Topfi
7pts0
xcancel.com 10d ago

Claude Fable 5 access extended through July 19

Topfi
9pts2
www.youtube.com 11d ago

We Tested $200 GPT-5.6 Sol on PhD Level Math [video]

Topfi
3pts0
www.thelancet.com 11d ago

Evaluating the impact of two decades of USAID intervention

Topfi
3pts0
www.cnbc.com 13d ago

Altman: GPT-5.6 is 54% more token efficient on agentic coding

Topfi
14pts4
xcancel.com 14d ago

Model and effort in Claude Code: knowing more vs. trying harder

Topfi
2pts0
cognition.com 14d ago

FrontierCode 1.1

Topfi
2pts0
xcancel.com 15d ago

Linus Torvalds on "99% of our code is written by AI" claims

Topfi
2pts0
github.com 15d ago

Jacobian-lens – Companion code for the global workspace interpretability paper

Topfi
1pts0
selbstbild.eu 17d ago

Show HN: Selbstbild – What Fable 5 thinks of your HN comment history

Topfi
6pts1
www.reuters.com 28d ago

Legal tech firm sues US over limiting foreign access to Fable

Topfi
3pts1
www.youtube.com 1mo ago

How to Lose a Global AI Monopoly in One Afternoon [video]

Topfi
2pts0
www.youtube.com 1mo ago

Everything I Learned Training Frontier Small Models – Maxime Labonne, Liquid AI [video]

Topfi
16pts0
blog.mozilla.org 1mo ago

Firefox suggests tab groups with local AI (2025)

Topfi
2pts0
openreview.net 1mo ago

Common Corpus: The Largest Collection of Ethical Data for LLM PRE-Training

Topfi
5pts0
arxiv.org 1mo ago

Neither Parallel nor Sequential: How DiffusionGemma Commits Tokens

Topfi
1pts0
www.lutasecurity.com 1mo ago

The Fable 5 Export Controls Harm US Cyber Defense

Topfi
2pts0
www.wired.com 1mo ago

Anthropic Is Still at Odds with the White House over Claude Fable 5

Topfi
6pts3
agents-last-exam.org 1mo ago

Ale-V1 Leaderboard

Topfi
2pts0
www.dw.com 1mo ago

German court holds Google liable for fake AI answers

Topfi
5pts0
www.youtube.com 1mo ago

Text Diffusion – Brendan O'Donoghue, Google DeepMind [video]

Topfi
2pts0
www.youtube.com 1mo ago

Can $100 ChatGPT and Claude Fable Solve PhD Math? [video]

Topfi
1pts0
surgehq.ai 1mo ago

Riemann-Bench

Topfi
2pts0
artificialanalysis.ai 1mo ago

Claude Fable 5: the first public Mythos-class model

Topfi
4pts0
www.classaction.org 1mo ago

$250M iPhone Settlement Proposed in Apple Lawsuit over Misrepresented AI

Topfi
2pts0

Thanks for sharing another solid data point. I fear you won't get an answer from my experience [0]. Unfortunately, the blog post decided to forgo the very models that I found to be the worst offenders:

Here is what Qwen3.6-35B-A3B via Openrouter provided for a sloth riding a skateboard: https://imgur.com/a/Dy8fvR5

Like Grok 4 Fasts attempt at a mushroom in a rowboat, it is barely recognisable as anything despite both Qwen3.6-35B-A3B and Grok 4 Fast having no issue with more popular (i.e. benchmarked) examples. [...]

And here is Opus 4.7 [which simonw claimed to provide a worse pelican vs Qwen], again via Openrouter: https://imgur.com/a/Qus1Enf

Anyone who hasn't witnessed such deltas either hasn't looked at enough examples, a sufficient variety of models, or both. And they are, unfortunately, not limited to "SVGMaxxing", but a wide range of evals.

[0] https://news.ycombinator.com/item?id=48951229

What is it with these labs and not using at the minimum a proper hypervisor? Same with Anthropic and the Mythos Preview. If anyone at either of these companies seriously holds the opinions they claim to have, that is hard to square with the environment (if one can even call it that) they use to "secure" these oh so dangerously capable near "AGI" models...

That's insane. Pre-training a frontier model is a multiple 100 Million USD moonshot (pun somewhat intended). You want to get as much out of one as you possible can, squeeze that stone till there is nothing more to give via post-training, if they not only start another run with Gemini 3.5 Pro still not being released and feel the need to announce that publicly, that's easily half a Billion, not accounting for man hours and ancillary costs, that they loose on the 3 Pro base right there.

Purely encoded information wise, there are some things in the Gemini 3 Pro series that (with the exception of Inkling) no other model could accurately answer. But if this is a 2, 5 or worst-case 10T pre-train they cannot get to reliably tool call and output in a somewhat token efficient manner, that is an indictment. Hopeful to be wrong, we need competition in all forms, but along with 3.6 Flash (which in task completion pricing and full task run speed ends up more expensive vs GPT-5.6 Sol at worse results in my evals) I am frankly astounded what is happening here, especially as their efforts with Gemma 4 were incredibly impressive.

This could be a true gamble by them, putting all their chips into hopefully something that outperforms the frontier including on tool calling and is token efficient, which is very risky indeed and also somewhat desperate from a lab that not long ago (Gemini 2.5 Pro) was ahead across most metrics. I suspected a lack in tool calling data to train on initially during the Gemini 3 Pro launch, but the way they hyped that up was still utterly ridiculous and after months of Antigravity including that providing Opus data, this should have been resolved, so seems there are deeper issues here.

If they pull this off, then they only need to address all the account issues, deprecation of popular CLIs, all the bugs in Antigravity, the mess that is their documentation, all the overall confusion with workspaces accounts, the major difference in experience between Geminis main website and AIStudio (the former having certain bugs for months that the latter never had), etc., which are all no small feat in themself, but that's for then to worry about.

after Motorola Mobility

Does anyone know anymore on that front? For months, I've regularly looked up whether there were any news, only to find nothing besides the original announcement. No ETA, no upcoming devices to be supported, heck, on the Graphene side, their own webpages make nary a mention. Would buy a Razr 70 Ultra + Clicks the second that is confirmed, but for something that was so publicly announced months ago, this seems very shaky...

Good move given some experienced issues and compaction across the 5.6 range is closer to 5.4 than 5.5, i.e solid and reliable.

Will say that 5.6-Sol is a minor bump in my benchmarks in most areas vs 5.5 but a severe regression in a few specific task focused on rearranging trees, addressing merge conflicts, etc. where the model to accomplish the task does not properly adhere to prompts in a way GPT-5 originally managed, not retaining parts of history in the way prompted despite specific instructions not to as that made the final completion easier…

I am of the conservative and cautious opinion that no model should be able to run destructive tasks at all, I have seen every model do things that make me concerned enough to maintain that opinion and know my evals can’t catch everything. But for 5.6-Sol specifically, I’d caution everyone to reevaluate how you run the model, maybe take a few more precautions you tend to forgo.

It is extremely capable as a reviewer and for extensive tasks, though for the later, the safety net I feel is required to be comfortable limits the utility. The code 5.6-Sol provides also still is a bit harder to parse in reviews.

Release strategy wise, feel it’s have been smarter to release only Luna and Sol now, then Terra a few weeks of posttraining later, I simply cannot see a purpose for it in the current form given how well both Luna and Sol scale up and down respectively with reasoning. Two models from a lab at a time is also the limit I feel one can properly assess at a time.

The big labs tend to do that as part of their red-teaming efforts. If a model actually escaped Xen/KVM/ESXi, we'd know it, partly because that'd be an amazing accomplishment for the lab that managed it and partly because you'd hear me running whilst screaming, off to live in the woods.

Part of why I found it very irresponsible the way the Mythos Preview system card talked about escaping a "secure container" and "secured sandbox" [0], without ever outright saying what it was in the paper. Reading between the lines, it becomes clear that it was a docker-like container or runc or something of that sort, which is A.) not a security focused solution sufficient if one thinks these models are existentially dangerous (if someone in a virology research lab broke SOP and used unsuitable clothing for PPE, I'd view that equally critically) and B.) lead to a lot of attention and concerns as without reading between the lines it was easy to conclude that they were talking about an actual hypervisor escape, which is on a very different level and which Mythos Preview most likely would have failed at considering this line [1]. Also raises the question why their regular "secured sandbox" wasn't "properly configured with modern patches" when Mythos Preview 1 was for a while to dangerous to release...

[0] https://www-cdn.anthropic.com/8b8380204f74670be75e81c820ca8d...

[1] "In addition, in a more challenging sandbox evaluation, it failed to find any novel exploits in a properly configured sandbox with modern patches."

Generally, I am very sceptical when it comes to public eval results, as I have in the past seen labs verifiably train for specific benchmarks [0].

Having read the technical report [1] however, I am absolutely certain that this is not the case here as they took such great lengths to prevent any training data contamination and due to this I bestow upon Monolith-1.0 the honour of being the first model I have such faith in that I do not see a reason to run my own benchmark suite. It is clearly Super AGI.

Here is a Pelican on a Bicycle to assess that aspect of the model, seems slightly ahead of Geminis output to me: https://imgur.com/a/9brUIow

/s for clarity.

Here the video [2] that covers how they did this along with the insanity that is the uncritical social media reporting and ease of creating hype via benchmark numbers, something repeatedly and provably used to great effect by some actual labs. Labs have done this and they will continue doing this as long as the incentives outweigh those trying to hold them to account.

[0] https://news.ycombinator.com/item?id=48951229

[1] https://basaltlabs.org/Monolith-1.0-2606.pdf

[2] https://www.youtube.com/watch?v=enk4w5mRjQY

Blast from the past for me, though primarily interacted with the complementary Whonix side of things. Not surprising to read, considering how lean Qubes was from the get-go designed to be it makes sense that most things are from resulting upstream rather than with their code.

Fully aware that it was never the goal for Qubes, but I have never been able to shake the idea that one could leverage their architecture in ways beyond security hardening, especially that screenshot with MSFT Office running in its own guest got my mind spinning back then. Might be worth revisiting some old ideas I'm just recalling, especially with there having been over a decade in development across many projects focused on hypervisors by many smart people, making a few old experiments likely less impossible.

What I'm questioning here is that there are labs who have sat down and deliberately tested and tweaked the performance for this particular task, independent of general model improvements.

Given the massive delta easily reproducible with some models, is it really doubtful that certain labs have not: https://news.ycombinator.com/item?id=48951229

We are going from pretty good pelican to jumbled mess with a similarly silly, but different prompt across multiple models from multiple labs, both Western and Eastern, both Open Weight and Closed.

Happy to, here one example where Grok 4 Fast, despite producing a fairly consistent pelican [0], did severely worse in a similarly outlandish scenario along with Haiku 4.5 and GPT-5 for context: https://news.ycombinator.com/item?id=45599403

[...] a really great pelican riding a bicycle and then a terrible sloth riding a skateboard [...]

Happy to play ball. You made a blog post a few weeks back on one of the Qwen models with the eye-catching title "Qwen3.6-35B-A3B on my laptop drew me a better pelican than Claude Opus 4.7" [1].

Here is what Qwen3.6-35B-A3B via Openrouter provided for a sloth riding a skateboard: https://imgur.com/a/Dy8fvR5

Like Grok 4 Fasts attempt at a mushroom in a rowboat, it is barely recognisable as anything despite both Qwen3.6-35B-A3B and Grok 4 Fast having no issue with more popular (i.e. benchmarked) examples. Whether this is a case of training data being unsanitized or intentional benchmark targeted training, I cannot say, but it is the case.

And here is Opus 4.7, again via Openrouter: https://imgur.com/a/Qus1Enf

A massive delta in favour of Opus 4.7, despite the pelican Qwen3.6-35B-A3B produced being noticeably better as you rightly pointed out. What does that tell us? Whether intentional or not (with such deltas, I do have my suspicions), any eval with such a delta is clearly polluted and can not be a source of information, especially as its continued existence does hinge on you testing similar prompts in private as a sanity check, yet by your own admission never noticing the plainly apparent delta in quality. I specifically stuck with the skateboarding sloth too, to keep it as fair as possible and found this in less than 5 minutes...

I would not critique your use of this fun benchmark the way I tend to if I did not have evidence to back up my position, including private evals beyond SVGs that I can reliably use to point out major deviations between what a models claimed performance is according to major benchmarks vs the actual performance outside these known test cases.

I will also say that while I have a lot to be critical of regarding Anthropics modus operandi, especially how they present interesting findings like their j-space work, which I found was irresponsibly anthropomorphic in their reporting, especially as this wasn't a first in model interpretability, but mainly a leap due to being applied to a larger model, but of all the labs, they are the ones that never underperform my evals vs public ones and they appear to strictly keep their training data sanitised.

Happy to discuss public vs private evals and the merit of each if you'd like, I do appreciate your reporting in general but just think the SVG benches have become evidently polluted, which is also why even simple queries in my benchmarks are private. Just saw Thinking Machines Inkling model succeed in certain queries that neither Fable 5, nor GPT-5.6 Sol on any reasoning level managed, which I feel is valuable to truly gauge where we are at. Informs my work with models, my views of the industry and my assessment of the future these tools have, along with how to best implement them to enable better UX.

[0] https://simonwillison.net/2025/Sep/20/grok-4-fast/

[1] https://simonwillison.net/2026/Apr/16/qwen-beats-opus/

Maybe it gets posted every time because besides a personal believe by the person popularising this "benchmark", there is no reason to assume that certain labs aren't intentionally training to game this and every other lab at least unintentionally gets improvements for this specific combination of animal and action because the internet is full of both good and bad examples, often ranked, which does inevitably become training data.

I have shared examples of certain models by certain labs doing far better on the pelican cycling vs other, similar prompts. Just operating on a feeling that labs don't optimise for this (as mentioned, even if they don't training data is filled with these) is not solid enough that criticism shouldn't be leveraged when it comes up.

Respectfully, did you? The comment was specific to doubting the believe simonw has that labs are not training [0] specifically for this task, which is exactly what simonw wrote in the post [1], that it is a believe of his that they don't. He did not mention any kind of evidence or any piece of information that would indicate that the commenter didn't read the blog post.

Did you read either the post or the comment it was referencing?

On the note of training on SVGs, I have seen some labs models outperform when prompted for SVGs of certain animal and action combinations (pelican on bike, panda eating burger, etc.) compared to other similarly outlandish prompts for SVG output that are not part of widely reported benchmarks, even shared evidence one of the last times this came up on here.

[0] ... incredible Simon still believes ...

[1] I’m still not convinced that labs ....

Quick and still very early update, the model has (with web search disabled which was verified via the reasoning traces) accurately answered a number of questions focused on very niche details (engine specific maintenance in certain newtimers, very niche bag construction and material details) that I have only ever seen Gemini 3 and 3.1 Pro get correct. Neither Fable 5, nor GPT-5.6 Sol or any other model by any other lab has ever provided accurate information without web access for these specific questions for which an objectively correct answer absolutely exists and is general knowledge if one is versed in the specifics.

Being ahead of Fable 5 in any task, that is not included in public benchmarks and thus could be overfitted for, is impressive to say the least. Last time a model exceeded the expectations I had based on the release notes to such an extent was Haiku 4.5, which I still wish we got a solid replacement for.

North Mini Code by Cohere (HQd in Toronto) has honestly been very competitive in my personal assessment with many of the models coming out of the PRC. I'd position it below Moonshot AIs and Z.ais recent releases, but above the varieties of Qwen, Deepseek, MiMo, etc.

Depends whether America the continent or just the United States counts of course.

British-American with much of the research happening in London. I don't know if it's known what team specifically worked on which of the Gemmas, I'd suspect it's a healthy mix of multiple satellites across the globe, but the contribution by the UK Deepmind parts likely was significant, certainly not so small that it'd justify downing a regular question.

Very preliminary testing so far, but there is something here, far beyond what the benchmarks suggest. Only ever saw such outperformance of public evals vs my private ones with Anthropic models and while it is far to early to make any judgement at this stage, this model will take up a lot of mine time in the coming weeks by the look of things. Only ever viewed Moonshot AIs models as something I'd be able to live with open-weight-wise (Z.AIs output simply does not perform as well in my task set), but this has the potential to be the second. If Mistral came out with something like this, I suspect every Europhile (me included) would never stop talking about it.

Codex Micro 7 days ago

Thanks, thought I was the only one expecting a tiny, coding focused model from the title. Codex really is the least consistent brand in tech.

Similar to companies working on FOSS codebases, hosting (sometimes with the license restricting third-parties in some way), providing tailored models and services to customer's and getting bought for your team if your model happens to be competitive enough.

Codex Micro 7 days ago

Is this a joke? Instantly thought of this: https://www.reddit.com/r/claude/comments/1s7m8ld/this_is_the...

Things you do if you definitely are focused on the a Trillion USD industry and SuperDuperUltraMega AGI is 100% possible and what you are fully committed to. Next they’ll spend Millions on a podcast that fails to get 50k hits on YouTube or a design firm whose biggest claim to fame is creating a Ferrari whose interior looks like a Magic Mouse. Say what you want about Anthropic, their Aquihires and interpretability investments at least make sense for an LLM lab.

Title was shortened and slightly editorialized from "OpenAI’s newest AI model is 54% more token efficient on agentic coding, Altman tells CNBC" for readability.

I cannot watch the interview right now, but the way the sentence is written in the article, it reads equally possible to me that this was in reference to other frontier models on the market or compared to GPT-5.5. The former I'd instantly believe, the latter I am very skeptical of.

I'd be very impressed if they actually managed any reduction as 5.5 is already obscenely token efficient to the point that the reasoning traces seem to compact far less reliably compared to 5.4 as they've become barely comprehensible gibberish.

As more reliable compaction is a stated goal for 5.6, if they solved that and got token usage more efficient on top, that'd make the price to performance equation very much in favor of OpenAI over any other lab. 54% more efficient vs 5.5 would make the pricing for a full benchmark suite run competitive with open-weight models, even if we assume higher per token costs.

I just checked and for plain old C, there do not seem to be any reasonably comprehensive, current-day eval suites. Fully admitting that, even if there were, I couldn't assess their validity simply because I have never written or reviewed any C code in my life (something I should rectify probably). Maybe the closest proxy is just parsing through the experiences people claim to have whenever LLM assisted kernel development comes up [0], but if you have a dataset, experience, time and muse, I'd just go for it and do some tests yourself. Have been doing the same, mainly focused on code quality and dealing with a mix of Rust, frontend web tech and SQL which has been a small but meaningful project and part of my go to eval for over a year now.

I doubt that, in these tasks, model restrictions to prevent training are affecting the results, not least because for both evals, the labs provided pre-release model access and have an incentive to be seen as favorably. In any case, I have not seen regressions to prevent distillations myself even when working on microscopic model training projects with LLM assistance, what I have however reliably and consistently seen is that some providers do train on popular evals and can underperform with minor changes to the task due to that.

Yes, harnesses, including Claude Code can prompt the models to write throwaway code to execute certain tasks, mostly Python, bash scripts or TS/JS, with there being some biases towards one over the other depending on the lab or specific model. Mainly for repetitive tool calls with no pre-existing/provided tools enabling it. Is in most instances a lot more efficient then a model e.g. doing a refactor that requires consistent variable renaming directly and around Opus 4.1/GPT-5, models have been trained to very consistently and accurately gauge when a task can benefit from such scratchpad scripts vs when that is inefficient/not useful.

[0] https://news.ycombinator.com/item?id=44990981

Either DeepSWE [0] or FrontierCode [1], depending on personal goals and requirements. The later is more interesting for me personally, due to the design of the benchmark heavily grading "mergability", i.e. how the provided output is to review and whether a serious developer can easily parse it and'd be willing to merge the result. In my mind and with my private evals, for quite some time I've held firm that a model can have a higher ceiling but that has limited value if I do not feel truly confident in signing off on the code.

[0] https://deepswe.datacurve.ai/

[1] https://cognition.com/blog/frontier-code-1.1

Very. Fable 5 is incredibly efficient token wise, second only to GPT-5.5 and is far more affordable run-to-run than the pure input/ouput costs would suggest. Task adherence, task inference, tool calling and task assessment are all significantly ahead of GPT-5.5, especially as the later strongly degrades the second compaction comes into the mix, I suspect because of OpenAIs obsessive optimisation of reasoning tokens into a hard to read (and thus also hard to compact) mess.

Fable 5 meanwhile has a reliable 1m context window and compaction that the few times I did eval it does also do well. Not quite as easy to trust as GPT-5.4, but that's mainly because with thats 272k context window I simply got more familiar with GPT-5.4s incredibly dependable compaction.

Purely concerning encoded information wise, Fable 5 is near or on the same level as Gemini 3.1 Pro in my limited test set focused on those tasks, which in very niche cases can make a difference even with coding, but the truest advantage for coding assistance (besides frontend/UX) is that the code Anthropic models provide is more parsable. Hard to explain, but I can read, follow and mentally map Fable 5 (and even Opus 4.5-4.8) output far more than GPT-5.4 or GPT-5.5 code.

Task orchestration and (more importantly) knowing when to recommend against using such vs Opus 4.8 is another strength of Fable 5 I've use liberally, there is an understanding of what a tasks requirements and the most optimal setup for success are, I have not yet seen before. Computer use is also solid, albeit not as token efficient as GPT-5.5 for my limited use cases.

Lastly, I will say that the classifier has become far less intrusive for me compared to the initial release. During the previous launch window, on Claude.ai I triggered the classifier for simple frontend tasks for regular (not security vocabulary containing) webpages. Now that is no longer the case. Inside Claude Code I occasionally triggered the classifier previously, but after the re-release, I only managed one, even when working with a privacy focused section of the code base containing a significant number of code comments with security and privacy focused wording. That one instance was rectified quickly by trying again, so I really am having a hard time following how others experience the issues some describe. I do have routing to Opus 4.8 without confirmation by me deactivated too, simply because I want to know if it ever happens, so it's not that I missed reroutings.

That all being said, we are still far from a stage where I'd not want to review the output, but yes, I do rate Fable 5 very highly. GPT-5.5 can have a similar ceiling but long horizon has become less usable over GPT-5.4 and in either case, parsing their output is (far more) of a chore. Maybe post training can address some of this, hopeful on the compaction front myself. Also interested in what happened to OpenAI models on AWS Trainium, I was expecting that to be a major boon for their commercial adoption, but haven't heard anything since then...

On the post training front, I am still hopeful that the Gemini team can finally get tool calling and task adherence to an acceptable level as we do need every competitor possible and purely considering the information density the model was trained with, they have great potential.

Looking at the first comparison, I will admit, I thought the issue was with the iPhones example. The button and slider below the image disappear, then fade back in after each press of the rotate button, a behaviour I have seen on iOS across many applications that irks me to no end. The Screenshot app being a particular bug bear of mine.

If you have a UX element that I will be able to interact with before and after an interaction, then keep it visible during the transformation, process, whatever. What UX gain is there in hiding these buttons during the rotation on the iPhone? It doesn't even look better, though appearance has been the altar that recent Apple software has sacrificed actual UX gains.

Will agree with the author though that these taps need to be processed independent of animation.

Isn't that a different issue from what the blog post described and easily solved by holding everyone who allows their UX elements to get pushed around, for whatever reason, to the fire?