Apple isn't counting on their model to be a frontier coding and cowork model. Gemini is perfectly fine for the tasks that new Siri is supposed to be doing.
HN user
winstonp
That is not the Grok CLI being discussed. That's an open source, third party CLI. https://x.ai/cli is the official Grok CLI being discussed, and it is not open source.
The avg coding session has hundreds or thousands of tool calls. Even a 5% failure rate noticeably notches up token use and cost. See Gemini.
Doubt. Tried it, Grok 4.5 is leaps and bounds ahead of Gemini for real work.
And token costs vs. raw compute+electric cost is unit profitable too.
While I mostly agree with your post, I do want to point out one thing:
Or he could release a model trained largely by existing open weights models. Which without some huge breakthrough probably has no chance of surpassing them, so is pointless.
This seems to be categorically untrue. Composer 2.5 is a substantial improvement on its underlying Kimi base model.
same happened to Opus 4.7
the British are notoriously sensitive to heat. They'll call 30 Celsius weather a heat wave.
do we think he found it on craigslist
Gemini pretty clearly has the best underlying model, and the worst RL and post-training of the lot.
It is ridiculous to be building infrastructure with a seventy five year lifespan under an assumption that may never come in its entire life.
Car apps, beside Tesla, are universally awful to use. Even Tesla's is not beyond reproach (app size is massive, for one), but at least it doesn't make me want to poke my eyes out. Apple should make a "Cars" app that's like the "Watch" app and let them standardize.
They are absolutely clueless about how to talk to this administration.
CC TUI is often a culprit of memory leaks, image-pasting is a nightmare, particularly over SSH, flickering all over the UI, among others.
You can!
You just have to pay API prices.
google models are still very unreliable at actually calling the tools you want it to call.
I don't believe Anthropic trains on Trainium, only serves models on it.
Claude Code TUI is garbage. There's nothing worth protecting in there.
I do think Anthropic nailed their naming down much better than OpenAI
OpenAI's training is better suited to developing models that don't have these tendencies
I've never had a smelly Mac, and I've owned maybe 10 different ones across personal and various work laptops. And 90% of devs I've ever met have used Macs and none of them smell either, so it's zero out of maybe 200+ in my personal experience.
Google Search has been awful for the last couple of years. Good riddance to SEO.
Same. Overuse of the word "genuinely" is another tell.
Which apps have you seen ask for someone to setup a local LLM? Can't recall having ever seen one
The new Studio Displays have more powerful chips than the Neo.
Agree. Claude tends to produce better design, but from a system understanding and architecture perspective Codex is the far better model
Everybody says they are, with the main point being they can get the M6 Pros onto TSMC's 2nm node and save 3nm capacity for iPhones.
They are at least nice for comparing it with the max of the Intel. That should really say gives them up to 22 additional hours given the wear on their batteries lol
DeepSeek hasn't been SotA in at least 12 calendar months, which might as well be a decade in LLM years
This guy was into the bagscoin BS that was going on on X this past month. Wouldn't trust a word he says.
Once Chrome does, for many devs, you can simply enforce a version check and say "please use latest Chrome" and be done with it.