How much does it cost? I even made an account and I cannot find pricing anywhere...
HN user
resonious
So it's GLM-5.2 performance for almost twice the price.
That said, the speed looks really good. I think it's competitive with Fireworks's GLM 5.2 Fast, although Fireworks is still cheaper.
Why does KV cache matter if they show it's cheaper anyway?
The article shows that it's still cheaper to pay for the switch than it is to let the frontier model do all the work. At least in benchmarks.
If you want to interact with plans then I think this technique just isn't for you.
for frontier work.
I'll agree that GPT 5.6 may well be the best given the above contstraint, but for run-of-the-mill dev tasks (real ones, not benchmark ones), GLM 5.2 still blows every other model out of the water.
Cost per task as a metric is a bit ridiculous because there are so many types of tasks. GPT-5.6 can do some tasks GLM could only dream of, but GLM can do some tasks 100x cheaper and better than GPT-5.6.
Oh My Pi has it. I'm a big OMP shill right now. Seems not very popular, but it has the stability of Pi with the features Opencode (and more I think; OMP has web browsing and a more advanced edit system too). OMP often outperforms Claude Code and Codex for me.
Just features for the client. The same stuff I was shipping before but much faster and more stable.
Trust me there are plenty of us using cloud AI to actually ship stuff. We just aren't writing blog posts about it.
The artificialanalysis cost per task chart has DeepSeek as the clear winner and Fable as the clear loser. But I would still pick Fable for some tasks, so that also can't be all there is to it.
But I agree that price per token figure is not great. It seems even the tokens per character can vary between models, so it's basically useless.
But artificial intelligence, far more than any tool we’ve ever created, intends us not just to sit forward and behave, but to cease to think critically, to cease to imagine, and, most temptingly, to cease to feel struggle and pain.
If I don't think critically, AI causes me to feel struggle and pain.
I consider myself an LLM maximalist, at least in the context of software engineering. Even after delegating everything you can to LLMs, there's still work left for you. And it's far more interesting. Even the delegation to LLMs is interesting - it's not that easy to make it work well.
In my domain (web app CRUD), we build the same thing over and over and over. And it's not beautiful or interesting like furniture. It feels like with LLMs, all that repetitive stuff is gone and I can focus myself on the tricky problems.
Right I saw them saying something along the lines of "they're good at subagents". But this seems true even with third party harnesses. So I'm wondering what Codex is hiding.
I've also seen Google indexing pages with random values in the path that don't get linked to statically (server asks for the URL then redirects to it immediately). I'm pretty sure they index straight out of the Chrome address bar.
I guess this implies that non-Codex harnesses get a little bit worse? In wondering what's so special about their subagents system that they feel the need to hide these messages...
"This is the part where you think/learn" and "this is the part where you just let the LLM figure it out" is a fine format for advice nowadays I think.
I am getting a little tired of every single HN comment being about how the linked article is written by an LLM.
This seems like a lot. I told my oh-my-pi agent to build my ios app for me and send it to me remotely. After some chugging, it gave me a tailscale funnel link. And it worked. I didn't have to dig into any of this stuff. I didn't use the words "Developer ID-sign", "notarize", or "staple" in my prompt.
I think the moralities of all the big heads in AI are questionable. The training corpus is largely stolen, and they are all in inescapable debt but keep going. But at this point, their products are so useful that almost nobody is willing to sit back and wait for a "morally acceptable" LLM to come around (which would inevitably be inferior).
I can't comment on CSAM though - if X.ai really is "okay" with it then I'll agree with you that they're more immoral than the others.
~13 yoe, and I had some nasty WebRTC + CallKit problems that Opus couldn't make a dent on but Fable figured out.
This is an old rumor but I thought Nintendo made a loss on the devices. If that's true, why would they want to sell more?
Okay I hadn't heard of Vending-Bench until reading this and it was quite the ride learning about it through this article. Very fun read.
My very native programmer take is that it's not too surprising that their hacker model would be less ethical. The guardrails that separate Fable and Mythos probably wouldn't kick in during an environment like this.
Over optimizing and spending too much time on engineering decisions is also something an incompetent development team would do.
I think it's the tiny chance they it will help humans that makes it so fascinating.
Look, I felt it. I didn't wait for the official apology from Anthropic. I quite before they published that, then felt very vindicated when they did.
Deja Vu... This looks just like the Claude Code performance regression back in April. I just quit my Claude subscription when that happened and went to Codex.
Now I'm kinda thinking of trying per token for both, using GLM 5.2 on Fireworks for most tasks, shelling out to the big boys only when needed. Not totally confident I'll break even though.
Claude Code downgrades loudly but I'm not sure what happens over API or with other harnesses, OpenRouter, etc.
Yes you pay a big burst right after switching. After that, everything is cached and it's smooth sailing.
I have clients waiting for very gigantic features and the agent harnesses are a godsend.
This lines up with my experience with my mother, though it played out differently. In her case, she would switch doctors every ~5-10 years and each time they'd basically say everything the previous doctor said was wrong. First it was "you have Lupus", second it was "actually it's some other autoimmune disease", then it was "actually whatever you had has been in remission for some time now and you've been taking brain-numming medicine for no reason." Then it was "you have cancer", "it's a rare one", and "oh turns out the brain-numming meds have a correlation with rare cancers". The cancer part was handled well (albeit unsuccessful) though. After such a bad time with rheumatologists, I was shocked by how competent people were when it came to cancer.
All of the above was intertwined with brief stints with doctors that would just berate her for being a painkiller junkie, even though she hated the stuff and just wanted to find/fix the problem.
Kind of a rant, really. I'm not sure how to tie it back into AI. I do wish we had AI at the time so that we could at least cross-check, but I also understand that doctors are already sick of patients self-diagnosing on the web and that AI probably just makes that worse. At the same time, if our medical system could catch up a bit (more doctors? less corruption/paperwork? not sure what it needs) then maybe people would be less inclined to take matters into their own hands.
I live in Japan and yet can't seem to pay for their API in JPY... I bet their enterprise customers don't have that problem but it was pretty annoying given "AI in Japan" appears to be there only selling point.
It means scores well in common benchmarks.