HN user

cjsaltlake

57 karma
Posts1
Comments18
View on HN
GPT-5.5 3 months ago

Normal people definitely will not benefit from this by default.

SWE-bench was created to replace olympiad coding benchmarks. I think past olympiad coding benchmarks were much worse representative of real-world coding than something like SWE-bench, which is derived from real units of labor.

Further, olympiad style benchmarks are arguably easier to contaminate / memorize unless you refresh it regularly; but that goes for SWE-bench too.

If you read the mythos report, in which they discuss and account for contamination substantially, it still suggests that performance on SWE-bench verified is meaningful. Benchmarks, including SWE-bench can absolutely be gamed, but if you're not explicitly benchmaxxing, improving on SWE-bench still measures model improvements, at least up to the level of Mythos.

GPT-5.5 3 months ago

Labor replacing devices means nobody works in those fields anymore. If AI can do this for every field, nearly no one will need to work in any field. We'll have a giant fully automated resource-extraction machine.

Companies are entities with political interests like any other. The only way you can curtail their influence politically is through regulation. It is reasonable for entities to pursue their goals through any legal means.

India is a single country, with a relatively well established national identity (obviously not perfect!, but enough for people to work and live in the same area), and a relatively stable government. Most African countries are still in flux in a huge way, and are just barely starting to establish stable national identities after decades of civil war.