> This comes cross as unbelievable, and at worst a fantastical statement
You can read some of the recent reporting here: https://www.nytimes.com/2026/07/09/business/china-russia-ai-...
HN user
> This comes cross as unbelievable, and at worst a fantastical statement
You can read some of the recent reporting here: https://www.nytimes.com/2026/07/09/business/china-russia-ai-...
Yeah, I'm not sure what the purpose of this article is. I don't have a lot of experience with Rust, but I think this is the type of task even Opus 4.5 could do.
> I don't think people realize how much dumber the models get at larger contexts and how much more the token cost is.
It does not match my experience that the model gets significantly dumber. It does get slower and more expensive, yes, but that's a sacrifice that needs to be made when working on anything complex.
My process involves having the main agent use subagents to explore what is needed for the given task. Then it writes a plan. Then it has the plan adversarially reviewed by more subagents and hardens it. After all is said and done, the 1M token window is 30-40% full. This flow would never work with 272k context, and in fact I've had to tone it down significantly for 5.6 Sol. Which, now that I think about it, probably explains why the results I get with it are inferior.
> They succeeded in spite of their tech choices.
It is rarely the case that technology choice is the make-or-break when it comes to whether a product is successful and achieves widespread adoption. Some choices are less ideal than others, but at the end of the day if you manage to make something that people want, the rest won't matter much.
Ah, so you aren’t a lawyer then.
I take it you are a lawyer specializing in NY real estate law, then? Would be interesting to hear a more detailed analysis if so.
They can do it often because they have compute, and they have compute mostly because they are pretty far behind Anthropic when it comes to heavy enterprise users. They’ve supposedly added several million since 5.6 launch, which is a steep growth curve. We will see if those users stick around. Anecdotally, I’ve stopped using it for anything major because it has attempted to do some very unsafe things when I wasn’t looking. Friends I’ve talked to have also gotten over the initial honeymoon period.
I've used GPT 5.6 Sol Xhigh extensively since its launch, alongside Fable 5.
My impression is that it is about as intelligent as 5.5, but they dialed up the relentlessness meter to eleven. This makes it more likely that it will accomplish the task you give it, which I think is the primary reason it looks competitive in benchmarks. However, it also makes it more likely that it will resort to... unconventional, weird or outright unsafe methods to do it. So I have to watch it like a hawk.
The other day it tried to read env variables from prod using a CLI command. The task it was working on did not necessitate doing that even remotely. I have the SSH keys for that particular CLI tool tied to my 1Password. So when the agent failed (because I never authenticated the SSH key access), it wanted to take over the computer, for which I got an OS prompt. At that point I stopped the agent and asked it why it did that. It said it wanted to dig around 1Password itself to see if it could get the key. I asked it why it needed prod env variables, and it thought for a bit and admitted it actually shouldn't. So as of yesterday I stopped using the "approve for me" mode and now use it only for simpler tweaks and bug fixes.
Fable is not only more intelligent, but also way more insightful. It can sniff out my intent far more effectively, and its "real world" knowledge allows it to act as a seasoned product manager with domain expertise. It can also think outside the box and make suggestions that I would not have thought of. With GPT 5.6 I have to be way more literal.
I think this is a good feature, but should be gated behind a toggle that is off by default, and designed to be enabled per session via prompt.
There are situations when I want Claude to start working on something just as I'm about to head to bed or otherwise step away. It's kind of annoying to come back only to find that Claude worked for just 5 minutes and then decided to pause and ask a question.
That said, I think certain types of questions should not be automatable. Maybe it's already built that way, but I wouldn't want Claude to go with its recommended direction for anything related to operations like deletions, changing external systems, etc. Basically, things that cannot be undone should be a hard-block and wait for user input always.
> I pretty sure OpenAI and Anthropic are doing the same or worse.
So in your opinion, they are training on your data even if you toggle the "don't train on my data" checkbox off?
That's a bold assertion.
> Or fear of standing out?
Absolute irony coming from someone who uses a Thinkpad. ;-)
Yes, this is the right question to ask. A lot of the decrease in emissions in Western nations is the equivalent of dumping your trash in a lot across town. Sure it's "gone", but only in a technical sense.
We are already living through a mass extinction event: https://www.worldwildlife.org/news/press-releases/catastroph...
+4C would be game over.
> Oh, and of course, the wonderful world of hacked clients. The idea of anarchy was gripping!
My multiplayer experience in Minecraft consists of spending three days building myself a fancy treehouse, then logging in the next day and finding that someone had used an exploit to turn all of it to lava.
I guess some people do like chaos, but not me.
> El Paso has "major" connections to New Mexico.
El Paso is not part of the Texas grid (ERCOT):
https://www.epa.gov/green-power-markets/us-grid-regions
Neither are the panhandle nor the northeastern portion.
About a decade ago I worked with a product manager who used that phrasing constantly, so it kind of stuck with me.
There's parts I enjoy, parts I dislike and parts I hate with a passion.
However, I love putting something in front of users, seeing them use it and get value out of it. And AI lets me get there 10x faster.
> There's more to it than that: writing is thinking. If you stop writing code, you aren't thinking anymore.
Humans have been thinking long before writing was invented. Why is code special?
I've been climbing for a decade, but over the past 3 years I've put on a bunch of weight due to work and certain life events. But I want to change that.
I know what motivates me: seeing progress. The feedback loop of "do X, see Y gain" is what keeps me going.
So I started building an integrated dashboard that can aggregate data from multiple systems:
- My digital scale
- Apple Watch (sleep + running performance)
- Beastmaker Motherboard, which is an electronic board that you attach a hangboard to and it shows you various stats like how much force you're applying
The idea is that every morning I'll open the dashboard and be able to see exactly how much progress I've made the previous day: weight loss, strength gain, cardio performance.
It's an interesting problem. There's essentially two parts to it: Apple Health, which aggregates data from the scale and the Apple Watch and can POST-export it hourly, and the electronic board, which sends data via BLE in real time. The destination for both of these will probably be an always-on Raspberry Pi 5, but I haven't decided yet. Then I'll have a small server app that can pull the data from the Pi and draw some fancy charts.
> OpenAI also has infinite money
Except OpenAI needs every cent of that money for compute, and they don't have healthy profits that can replenish what they spend.
Their financial situation is simply not comparable to that of Apple's.
> I think the article kind of gives you the energy you enter the reading with.
No, not really. I know nothing about either of the people involved, and after reading the article I came away with an extremely negative opinion of the author. They are using their position as the leader of a language to attack and denigrate someone who used that language for years, donated to it, and eventually decided to move on.
Yes, this is exactly why I said yesterday that OpenAI does benchmaxxing, seemingly quite a bit [1]. I got a flurry of downvotes for it at first, but I think people came around to it once they tried the model like you did.
Ultimately I think the issue is that OpenAI is under tremendous pressure to perform, but GPT-6 is not ready yet, so they had to push GPT-5 to its limits, and the only way they could do it was with really heavy RLHF, which has its shortcomings. Like, it is super obvious that Sol, Terra and Luna are all heavily biased towards working on a problem relentlessly because that's what their reward functions emphasized. That pushes up their scores in some benchmarks but does not translate to actual intelligence and capability.
My Codex app got upgraded to the new unified ChatGPT app. I don't see Sol available though. Only Terra and Luna. I'm on the Pro plan. Anyone else see it?
The charts are also extremely difficult to parse. They seem auto-generated. Dataset coloring is atrocious.
Regarding your main point, yes, I agree. My impression (as someone who uses both Codex and Claude Code daily) is that OpenAI does a fair amount of benchmaxxing.
The second paragraph has four mentions of Fable. I think that makes my case pretty clearly.
CTRL-F: Fable
15 hits
Holy shit. They must be feeling very threatened by Fable if they're spending this much energy talking about it in the release notes for their own model.
UI: SwiftUI (primary), UIKit (limited), PhotosUI, WebKit, MapKit, Charts
Data/Concurrency: Foundation, Combine, Observation (@Observable)
Platform: CoreLocation, UniformTypeIdentifiers
Tests: Testing (Swift Testing)
FrontierBench
The problem is that the remaining 10% can bite you in bad ways.
I was in Cotswolds, UK a couple of months ago. For those of you who don't know, it's a rural region known for its "chocolate-box" villages and honey-colored limestone architecture. Basically, you go from village to village, most commonly via bus, taking in the sights and doing touristy stuff.
When planning the trip, my sister used ChatGPT, which helpfully (and relatively quickly) found the bus schedules and times for each hop.
Midway through the day, though, we ran into a huge problem: it turns out bus schedules are different on Sundays, and more limited. Which meant we couldn't actually go to our primary destination (the Model Village), and had to cut the trip short.
Yes, ChatGPT was quick and pleasant to use, but missed a crucial detail.
Afterwards I tried it with Opus and it did not make the same mistake.
It's also hilariously wrong. It essentially argues, implicitly, that those who don't communicate with other humans are missing out on the "most important thing in life" and cannot form a self-identity.