Bingo!
Yes the OP is completely biased.
HN user
Bingo!
Yes the OP is completely biased.
Most of those technologies being European, including the most difficult to make components in the system.
Why isn't there an American company doing this? Why isn't an American company making the lasers, mirrors, metrology, vacuum handling tech, etc?
The most absurdly biased and inaccurate take on ASML I've ever seen, congrats.
If the machines are so easy to build and integrate, why isn't it happening in America?
Humans make all these mistakes too, in fact, humans make more mistakes than AI does in coding at the moment.
Had a similar issue, several of my words weren't accepted.
I prefer 5.6 Sol to Fable personally and I've used both extensively.
Why should anyone put serious time and effort into using/understanding a product when the author hasn't put serious time and effort into making it?
I could make this exact same thing over a weekend and post it on Hackernews. But I won't because I would be embarrassed to do so.
The bar for posting something to HN should be high, the bar for wanting people to read your code, your writing, should be putting serious effort and thought into it. Not just vibe coding something up with a vibe coded README and 100% vibe coded code and not even a novel idea or implementation.
The AI generated README?
Well, you say that, but when "measuring" anything in RL, that measurement itself is not always obvious.
That is, creating the scoring system/judge models etc for RL is not easy at all. You can easily create an RL loop which is getting better and improving its scores, but actually the result is totally garbage, because you're measuring the wrong thing.
Can you explain how it works?
What problems would it do well on and why?
Where would it start to fail/break?
What are the limitations of a system like this?
When you vibe code a system in a complex area like RL, you basically have zero understanding of what its actually doing, whether its actually any good or not, what you're actually benchmarking, and when the system would fail.
It's the blind leading the blind.
I'm a guy who never cooks, I have no incentive to play dumb. I don't buy flour or sugar or other random raw ingredients in bulk. The closest I get is maybe a jar of coffee.
I'm from Europe, I never buy sugar, why would I? I don't want more sugar in my diet.
Do they? I don't recall ever seeing a bag of sugar in my life. I'm not a baker though so maybe that explains it.
A car is more easier to picture for me.
SWE Bench Pro is also gamed and shouldn't be trusted.
Probably means subjectively according to his own opinion...
How can this be "objective"? Surely its subjective.
I've tried a fuck load of harnesses but keep coming back to Codex as my harness.
The SWEBench benchmarks are really gamed at this point and should not be trusted period. The solutions are effectively in the training sets and have been for a while.
They are both excellent but excel in different areas. Fable is super super proactive and great for doing a LOT of work with a single prompt, also for creative work.
Codex is more details focused, often catches wonky bugs and correctness issues that Fable misses, feels more terse and less "friendly", more like a stern senior engineer versus a friendly talkative engineer (Claude). Codex is also better if you're already an engineer, Claude is better for non-engineers. I.e. Codex works better if you know exactly what you want and know the right way of explaining it.
"On Agents’ Last Exam (opens in a new window), an evaluation of long-running professional workflows across 55 fields, GPT‑5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive reasoning) by 13.1 points. Even at medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. That efficiency extends to smaller models, which are essential to making intelligence more abundant and affordable: GPT‑5.6 Terra and GPT‑5.6 Luna outperform Fable 5 at around one-sixteenth the cost. "
Some pretty big claims and results! Excited to see how it feels during usage.
I use Fable and 5.5 extensively and I still find both have a place in my toolkit, i.e. Fable IS good but it isn't perfect, and it's still better to play them off against each other. I have Fable and 5.5 write plans and have them adversarially review each other's plans.
Having this amount of competition in the coding model space is good for all of us.
I mean, I literally asked about the effects of nicotine withdrawal on the body after quitting smoking and the model got downgraded to Opus...
So yeah, if I can't ask about nicotine withdrawal, then I think almost anything biology related is going to get downgraded...
I asked fable about the effects of nicotine in the body when quitting smoking and got downgraded to opus.
Read the latest semianalysis article.
Anthropic is definitely profitable now, in fact, they’re crushing it.
Dude it’s not trivial to switch because the behaviors are different!
You’re clearly not building a product based on an LLM.
I’m still using various old Anthropic and OpenAI models for products I’ve built and released because I can’t risk the behavior changing in unpredictable ways and the users being pissed.
It’s much easier to switch out some deterministic software than an LLM which you’ve spent a ton of time on testing and benchmarking and understanding its nuances. Changing it is like replacing an employee who’s critical to the business.
I use both in different scenarios
For me it's the exact opposite, Anthropics models seem great for "vibe coding" by non engineers. My girlfriend uses Claude and loves it because she doesn't know any of the terminology and Claude happily fills in the gaps.
For me, with 20 years experience engineering across the stack for venture backed companies to FAANG, I cannot handle Claude at all, it writes way too much garbage that I never asked for. Codex is like a surgical instrument, it does exactly what I want it to and never bloats the codebase.
Anyone spending days with Claude with almost inevitably end up with a bloated buggy mess. Note: Codex also finds bugs and correctness issues that Claude misses, again, I've seen this probably 90% of the time. That is, Claude will happily tell you the feature is complete, but then get Codex to review the code and it will find 2-5 actual correctness bugs. Take those bugs and give back to Claude and it will admit it fucked up.
I've seen this behavior again and again and again. If you're not a strong/experienced engineer, Claude can seem perfect, but it's writing buggy code and you're just not aware of it, unless you're double checking with Codex or another LLM.
MCP is a protocol, that's all.
Saying MCPs are the new websites is like saying "SOAP" is the new websites, or "REST" is the new websites.
MCP is basically the AI equivalent of a REST API, it's not a product anymore than JSON is a product or XML is a product.
MCP makes token use WORSE, not better.
It's interesting how you would say this about China but not about the US, especially given what's happened recently with Anthropic and the US govt.
Do you really think the US government doesn't get access or couldn't get access to any of your chats with Claude?
"Brown people existing doesn't hurt you"
Yes, existing doesn't hurt. But when you import mass amounts of people who don't talk your countries language, have no intention of learning, and have no intention of getting a job, and on top of it are intensely religious supporting a religion which is antithetical to your values, then it DOES hurt.
Western Europe has simply allowed too many Muslim immigrants in than they could ever successfully have integrated. Now Europe is full of ghettos of Muslim immigrants who don't get jobs, don't learn the languages, and sap resources which the countries cannot afford.
Note: this isn't about being racist or not, it's just common sense. There is a limit of how many immigrants a country can take, financially and culturally.
Why was Uber valued in billions for years while making zero profit?
Why was Amazon valued at billions while making zero profit?
The stock market prices companies by many factors, revenue and profit are factors but so is growth.
Utilities companies make lots of profits but they are valued badly because they don’t grow at all!
Markets are forward looking and space is seen as a huge growth driver for the future, also RocketLab has been growing their top line revenue massively over the last few years.