HN user

meowface

12,008 karma
Posts13
Comments4,916
View on HN

This would seem to lead to absurd implications. What if in, say, 20 years, an AI is able to independently prove nearly everything important in under an hour, including the Riemann hypothesis, with no contamination from other proofs? (Let's say it's also free to write and run arbitrary code, as well.)

Whether it's 5, 10, 20, 50 years, obviously the takeaway cannot be "something had gone terribly wrong". The takeaway would be the smartest humans were never close to the theoretical intelligence and wisdom ceiling and never could've been. This will one day seem obvious in retrospect. There's no reason evolution by natural selection would've landed any species near such a ceiling.

You can use that retroactive logic about any hard problem though. Unsolved murder cases, math, theoretical physics.

If tons of smart humans try for years and fail and then an LLM tries for a few weeks or hours and succeeds, the implications are clear. And these are by far the dumbest LLMs will ever be.

I don't know why you're saying this, given this is not sycophancy and instead is an actual example of Claude Fable finding a real counterexample. I would get it if the mathematician had in fact posted something untrue or crankish, but he posted something true.

Frontier models have solved several major open problems in mathematics in the past few months, so this should not be a huge shock. "Anti-AI psychosis" will probably grow to outcompete AI psychosis by year's end.

The interesting thing about using Claude Fable 5 is it's nearly as irritatingly sycophantic as past Claudes while genuinely being smarter than the previous models. So you get a kind of yo-yoing of it glazing you as a creative genius and disappointedly revealing to you that your ideas are bad and dumb.

Oh yeah, I've noticed the same, for sure. But that's basically what I was expecting. The best model is not necessarily always going to be the most cost-effective/sensible model for a particular use case.

Fable-class models will probably be cheaper for Anthropic to serve within the year, though. And rumors are GPT-6 is of similar size and intelligence to Fable and may come out within the next few months. OpenAI models tend to give you more bang for your buck, probably in part due to OpenAI being able to throw more capital and compute around on top of being particularly willing to loss-lead to stay competitive with Anthropic.

At this point I barely put any value in any of the benchmarks. I just use the models for coding (and related things like software product design/planning/ideation/etc.) tasks and judge them subjectively, and also see how others judge them subjectively on HN and Twitter.

- I agree that trying to hamfist a poorly-engineered game-like rendering engine into Claude Code was not wise - hence my agreement in my initial post that "Claude Code is seemingly currently not very well-engineered".

- All of the stuff unrelated to rendering/display is the complicated stuff I was referring to. It's actually not easy to get an agentic harness (even one with, say, the simplest TUI imaginable) to work as well as Claude Code and Codex do.

- I included "necessarily has to be" to separate the two - it does not necessarily need to have a weird buggy rendering engine thing, it does necessarily need to have lots of careful agentic massaging.

- I do not understand how the implication "writing a game engine is easier than writing Claude Code" could have been drawn from my reply.

I promise your post was already downvoted when I saw it. It is possible some people upvoted it afterwards, changing the net karma.

It is possible you were not intentionally choosing to use an LLM to write/modify your posts, but they largely read like LLM output. The tool you're using may use an LLM and may be rewriting significant portions of your text.

But just to be clear: you are absolutely, completely denying an LLM played any role in writing any of your comments on HN? You are certain? You're claiming you did not use an LLM in any way in the production of them? Not even to "edit"?

And for the record you were downvoted by other people long before I saw your reply.

I'm not sure what you're saying. They spent ages adding guardrails to Mythos. Then they spent ages creating a whole new even more guardrailed version of Mythos called Fable. Then they added tons of classifiers so API requests to Fable would get rejected even if you ask a question like "what is a molecule". They put the thickest layer of bubble wrap around the model of any model in history. And then just today they made the classifiers even much more extreme than at the initial launch.

If they were truly honest in their beliefs of the potential risks of this model, how would their behavior have differed? I would expect exactly the behavior we see, if they were being honest in their belief.

Also note Dario here saying they shot themselves in the foot commercially with how they handled the rollout of the model - you can tell by his reflexive reaction how ridiculous he considers the accusation: https://youtu.be/v1wZwxY3CMg?t=2103

I can't roll my eyes hard enough at all the people who say this shit about Anthropic every day. I know I'll get downvoted. I know it's lame to complain about future downvotes. I don't care anymore.

Anthropic was correct in their assessment and early warning of Mythos's capabilities, and they did this rollout pretty well. They were not hype marketing. They were being genuinely cautious and honest.

The Trump admin was largely unreasonable with the sudden export control. (Though not entirely unreasonable.) The export control also had not much to do with Anthropic's pre-release warnings. See: GPT-5.6 currently being held up by the federal government.

Yes, I am pretty sure it was simply poorly worded.

They almost definitely mean "you will notice even more false positives during seemingly routine coding/debugging tasks than you did at the initial launch". Which is not surprising, given the ordeal they've been put through. Hopefully it won't be too bad.

The main depressing thing for me is it's now only 7 days on the subscription, and then full API pricing, with no mention of even a plan to bring it back to the subscription in the future. (The initial launch mentioned two weeks of subscription, then API pricing, then a hope to return it back to the subscription not long after.)