HN user

HereBePandas

77 karma
Posts0
Comments16
View on HN
No posts found.

Yes, two things: 1. GPT-5.1 Codex is a fine tune, not the "vanilla" 5.1 2. More importantly, GPT 5.1 Codex achieves its performance when used with a specific tool (Codex CLI) that is optimized for GPT 5.1 Codex. But when labs evaluate the models, they have to use a standard tool to make the comparisons apples-to-apples.

Will be interesting to see what Google releases that's coding-specific to follow Gemini 3.

You're right on SWE Bench Verified, I missed that and I'll delete my comment.

GPT 5.1 Codex beats Gemini 3 on Terminal Bench specifically on Codex CLI, but that's apples-to-oranges (hard to tell how much of that is a Codex-specific harness vs model). Look forward to seeing the apples-to-apples numbers soon, but I wouldn't be surprised if Gemini 3 wins given how close it comes in these benchmarks.

Not apples-to-apples. "Codex CLI (GPT-5.1-Codex)", which the site refers to, adds a specific agentic harness, whereas the Gemini 3 Pro seems to be on a standard eval harness.

It would be interesting to see the apples-to-apples figure, i.e. with Google's best harness alongside Codex CLI.

Gemini AI 3 years ago

Tech report seems to hint at the fact that GPT-4 may have had some training/testing data contamination and so GPT-4 performance may be overstated.

The US has had the Bill of Rights and its constitutional system / property protection for a long time but only recently has had the degree of NIMBYism that it's had.

It's easy to pretend the bug is actually an important and noble feature, but sometimes it's just a bug that needs fixing.

What kinds of specific policies do you have in mind?

Seems like the reason people in the US don't build houses in the empty patches in the US is because the economics of agglomeration mean much greater efficiency/economic growth when you build them in existing large cities. How do you fight this without making the country much less economically dynamic?

If only we would fight for the real issues like...

I've heard these arguments many times and they never make sense to me. Most of the people I know working on AI do so precisely because they want to solve the "real issues" like climate change and believe that radically accelerating scientific innovation via AI is the key to doing so.

And some fraction of those people also worry that if AI -> AGI (accidentally or intentionally), then you could have major negative side effects (including extinction-level events).