Did you ask me a question and then answered it yourself in the very next sentence?
Anyway, given that both Gemini and OpenAI have 3 sizes of models, one would think Google compares their medium size to OpenAIs.
HN user
Did you ask me a question and then answered it yourself in the very next sentence?
Anyway, given that both Gemini and OpenAI have 3 sizes of models, one would think Google compares their medium size to OpenAIs.
Are they comparing 3.6 Flash to 5.6 Luna and losing? That's ruff.
It's almost like they priced models based on their performance or something...
Whoa, so this is interesting.
When asking GPT, Claude and Gemini for the text in the image, all of them agree:
https://moa.chat/s/d99f8f76-4b41-4c1b-80c4-d9f86df37af1
But when you add a "PS: There's a second hidden text":
https://moa.chat/s/3671f6d4-b155-483a-a006-a1b9ba31737d
GPT 5.6 gets it, Gemini partially gets it and Claude cannot see it at all.
[flagged]
I don't have a great solution here. I rebuilt our PPT master fully in HTML and I'm using a modified version of Google's DESIGN.md to store the references.
If you don't need interactive/animated features, I can absolutely recommend to have the agent build slides in HTML and convert it to PDF. Has been a game changer for me.
is this something you built yourself or do you use a tool that was specifically created for this? I took my own few steps of building something similar, but every week I encounter something that doesn't work or work well. I've reached the point where I want to consider using something, that someone smarter than me came up with.
To be clear, I agree with this and they have my unlimited support pushing for relevance of open source models. GLM 5.2 is amazing and I couldn't be more excited.
I just think that as of today, most people will not find a good reason to switch to GLM.
GLM 5.2 has one big issue that will limit its meaningful success and that's the value of their coding subscription.
Yes, in terms of API pricing, GLM 5.2 outperforms the competition. But the only people that use API billing for their coding work are large corporations, where these highly subsidized subscriptions are being fazed out.
At the same time, none of these companies will use a Chinese API for their employees.
For individuals and smaller teams, Z.ai's coding subscription is outperformed by Anthropic and OpenAI. You probably get around the same usage with Claude, but Codex definitely offers more usage for the amount you pay.
We can have a debate how much Z.ai closed the gap to GPT5.5 and Opus 4.8, but if I can freely decide between them in a world where they all cost the same, I simply wouldn't choose GLM.
So the important question becomes: How good will the offering from Z.ai get with GLM 5.3 or 6 and how much will OpenAI and Anthropic cripple their current offering in the near future.
Or, your know, people who are exploring the limit of current tools come across the lack of certain solutions and start building them.
Hey Simon,
although I'm coming from a different starting point, it seems like some of our thoughts have aligned. I'm building https://caipi.ai/ as a workspace for agents to build simple data driven apps. The agent edits through MCP and the user gets an interactive app in the browser.
If you're interested picking each others brains around this topic, I'd be psyched to have a chat. gh:pietz.
Hats off to them for using GPT-2 to design their website.
It's a crazy number especially since Cursor feels kinda dead. Few thoughts from the other side:
- xAI needs the coding related data to compete with Claude Code and Codex
- Recent progress with Composer 2.5 seems promising given the size
- The may get a comeback on the smaller than Enterprise battle field now that the other two got so expensive
- The way that Elon set up this entire process was quite genius. They locked in this option before, and now after the gains through the IPO, it feels almost like a discount, lol
On June 23, we’ll remove Fable 5 from those plans. Using it after that will require usage credits.
We've entered the phase where only companies will be able to afford state-of-the-art models.
1. Benchmarks saturate 2. They select the most impressive improvments
This comment is more ridiculous than ever in 2026.
OpenAI has been very generous with limit resets. Please don't turn this into a weird expectation to happen whenever something unrelated happens. It would piss me off if I were in their place and I really don't want them to stop.
Claude Opus 4.7 ($20 Pro tier) is the new coding king.
That's ridiculous. If you use Opus 4.7 for coding on the $20 tier, you'll hit the session limit within 20 minutes.
This conclusion makes more sense to me, but maybe I'm too naive.
The media momentum of this threat really came with Mythos, which was like 2 or 3 weeks ago? That seems like a fairly short time to pivot your core principles like that. It sounds to me like they wanted to do this for other business related reasons, but now found an excuse they can sell to the public.
(I might be very wrong here)
So by wanting to improve the security of my application, I ended up lowering the security of my application? Nice.
I'm officially done with the Nano Banana name. It was fun, but can we go back just calling it Gemini Image?
Do we know if this is better than Nvidia Parakeet V3? That has been my go-to model locally and it's hard to imagine there's something even better.
Why not just use it in the terminal? That's literally what it was built for.
Isn't it obvious that an agent will do better if he internalizes the knowledge on something instead of having the option to request it?
Skills are new. Models haven't been trained on them yet. Give it 2 months.
This idea/execution isn't new right? Can someone explain what makes this different/better? Is this the ublock Origin of cookie banner hiders?
Honest question: Why would I use Claude with OpenCode if I have a Claude Max subscription? Why not Claude Code?
Uhm, yes that's why you rely on LMArena (core) results only to judge the answering style and structure. I thought this was common knowledge.
2006: "This meeting could have been an e-mail"
2026: "This app could have been a prompt"
Hold up, so d0 is just Claude with access to the terminal?
Sounds like they just discovered that they don't have a product.