HN user

scellus

81 karma
Posts0
Comments55
View on HN
No posts found.

Perfect economic substitution in coding doesn't happen for a long time. Meanwhile, AI appears as an amplifier to the human and vice versa. That the work will change is scary, but the change also opens up possibilities, many of them now hard to imagine.

My work is better than it has been for decades. Now I can finally think and experiment instead of wasting my time on coding nitty-gritty detail, impossible to abstract. Last autumn was the game changer, basically Codex and later Opus 4.5; the latter is good with any decent scaffolding.

Thermal energy storage is one gotcha. It will eventually leak away, even if the CO2 stays in the container indefinitely, and then you have no energy to extract.

The 75% round-trip efficiency (for shorter time periods) quoted in other threads here is surprisingly high though.

It's complicated. Opus 4.5 is actually not that good at the 80% threshold but is above others at 50% threshold of completion. I read there's a single task around 16h that the model completed, and the broad CI comes from that.

METR currently simply runs out of tasks at 10-20h, and as a result you have a small N and lots of uncertainty there. (They fit a logistic to the discrete 0/1 results to get the thresholds you see in the graph.) They need new tasks, then we'll know better.

I find it odd that the post above is downvoted to grey, feels like some sort of latent war of viewpoints going on, like below some other AI posts. (Although these misvotes are usually fixed when the US wakes up.)

The point above is valid. I'd like to deconstruct the concept of intelligence even more. What humans are able to do is a relatively artificial collection of skills a physical and social organism needs. The so highly valued intelligence around math etc. is a corner case of those abilities.

There's no reason to think that human mathematical intelligence is unique by its structure, an isolated well-defined skill. Artificial systems are likely to be able to do much more, maybe not exactly the same peak ability, but adjacent ones, many of which will be superhuman and augmentative to what humans do. This will likely include "new math" in some sense too.

GPT-5.2-Codex 7 months ago

I like Opus 4.5 a lot, but a general comment on benchmarks: the number of subtasks or problems in each one is finite, and many of the benchmarks are saturating, so the effective number of problems at the frontier is even smaller. If you think of the generalizable capability of the model as a latent feature to be measured by benchmarks, we therefore have only rather noisy estimates. People read too much into small differences in numbers. It's best to aggregate across many, Epoch has their Capabilities Index, and Artificial Analysis is doing something similar, and probably others I don't know or remember.

And then there's the part of models that is hard to measure. Opus has some sort of HAL-like smoothness I don't see in other models, but meanwhile, I haven't tried gpt-5.2 for coding yet. (Neither Gemini 3 Pro; I'm not claiming superiority of Opus, just that something in practical usability is hard to measure.)

Opus 4.5 seems to be able to plan without asking, but I have used this pattern of "write a plan to an .md", review and maybe edit, and then execution, maybe in another thread,... I have used it with Codex and it works well.

Profilerating .md files need some attention though.

In other words, permanent instructions and context well presented in *.md, planning and review before execution, agentic loops with feedback, and a good model.

You can do this with any agentic harness, just plain prompting and "LLM management skills". I don't have Claude Code at work, but all this applies to Codex and GH Copilot agents as well.

And agreed, Opus 4.5 is next level.

No. In general the statistics look a bit amateurish, which is normal for a scientific paper. I'd actually like reanalyze the data, just out of curiosity. (Those p-values and other things can still be on the right ballpark even if the models and analyses are not top notch. I'm not exactly doubting them, and the results are interesting even without any correlation to UAP sightings or nukes.)

He writes as if only datacenters and network equipment remain after the AI bubble bursts. Like there won't be any AI models anymore, nothing left after the big training runs and trillion-dollar R&D, and no inference served.

Nostr 10 months ago

Same here. I like the idea, have tried the social-network side a couple of times, but my kind of content is missing or I can't find it.

https://bitchat.free now uses nostr for non-mesh contacts somehow, but I see no-one there either.

Even with pretraining, there's no limit or wall in raw performance, just diminishing returns in terms of the current applications, and business rationale to serve lighter models given the current infrastructure and pricing (and applications). Algorithmic efficiency of inference on a given performance level has also advanced a couple of OOMs since 2022 (for sure a major part of that is about model architecture and training methods).

And it seems research is bottlenecked by computation.

Yeah, just wanted to point out in my reply that the community is not paying the salaries their employees, the community in the sense of me and you, those who send observations and ids. The money mostly comes from big donors. To me that sounds a bit like bootstrapping the system, a startup if you wish. Moore giving them $10M per year from here to eternity is not a plausible future.

I'm not here to say it _should_ be open. Instead, I'm saying they offer a valuable, international (global) service and I want their economics to be sustainable, and have personally no objections to them keeping their AI models private if they wish so.

Meanwhile, the whole idea of iNaturalist has evolved around voluntary reporting, community involvement, and open data, and I think some of that needs to stay. They can't turn fully commercial.

Not currently, but imo they _could_ sell both data products and identification to research institutions. Like having the raw data still free, but charging for derivatives, or for professional service, incl. support, related to that data.

Especially selling identification services, which is related to keeping the models private, would make sense. Museums and various kinds of biodiversity monitoring schemes need mass identification, and having AI there to partially replace people would be a cost saving for the researchers and potential funding for iNaturalist. Offering such a service for free is neither practical nor justified.

(Meanwhile, I can imagine there to be lots of naturalist who hate the idea of their services being partially replaced by AI. It may lower the quality but the cost margin between a human and an iNat model is really wide.)

I think EU had a plan on using AI identification in some of their monitoring schemes. It could have been iNaturalist or someone else, anyway it demonstrates the need.

They get almost all of their money from Gordon and Betty Foundation, NSF, Nat Geo Society and the like. The "community" are almost entirely free-riders. I have donated a bit I think (maybe $20 or so), but their statistics say 0.24% of users donate.

That's probably not a sustainable situation.

I don't know. iNaturalist contributes a valuable identification service in their free apps. I'm a relatively experienced naturalist, but greatly helped by the iOS app, simply because remembering hundreds of species names is hard, even when the species are familiar. And the identification abilities of their models outside of my own domain are just stupendous, and freely available.

If they feel like keeping the models to themselves, I think it's a fair game. I give them observations, they gave me the id service for free. Maybe they even sell the models to fund their development efforts? I wouldn't mind... they need to fund their functions somehow anyway.

And remember, their observation databases are open. In fact my observations are automatically copied to the databases of a national biodiversity institution (which is open as well, except for some critical species).

Institutions need to maintain themselves and be able to pay their employees for them being able to feed their kids, etc.

So far it doesn't seem like winner-take-all, and all the major players (OpenAI, Anthropic, xAI, Google, Meta?) are backed by strong partnerships and a lot of capital. It is capital-intensive this round though, so the primary producers are big and few. As long as they compete, benefits mostly go to other parties (= society) through increased productivity.

No on AI, this is really a fringe environment of relatively uninformed commenters, compared to X. X has its craziness but you can curate your feeds by using lists. Here I can't choose who to follow.

And like said, the researchers themselves are on X, even Gary Marcus is there. ;)

It may be a talent drain too, but at least it's a selection bias. People just get enough and go away, or don't comment. At the extreme, that leads to a downward spiral in the epistemology of the site. Look at how AI fares in Bluesky.

As a partially separate issue, there are people trying to punish comments quoting AI by downvotes. You don't need to have a non-informative reply, just sourcing it to AI is enough. A random internet dude telling the same thing with less justification or detail is fine to them.

If one wants to follow AI development mostly in the sense of LLMs and associated frontier models, that's an excellent list with over half of the names familiar, to whom I have converged independently.

I have a list in X for AI; it's the best source of information overall on the subject, although some podcasts or RSS feeds directly from the long-form writers would be quite close. (If one is a researcher themselves, then of course it's a must to follow the paper feeds, not commentary or secondary references.)

I'd add https://epoch.ai to the list, on podcasts at least Dwarkesh Patel; on blogs Peter Wildeford (a superforecaster), @omarsar0 aka elvis from DAIR in X, also many researchers directly although some of them like roon or @tszzl are more entertaining than informative.

The point about polluted information environment resonates on me; in general but especially with AI. You get a very incomplete and strange understanding by following something like NYT who seem to concentrate more on politics than technology itself.

Of course there are adjacent areas of ML or AI where the sources would be completely different, say protein or genomics models, or weather models, or research on diffusion, image generation etc. The field is nowadays so large and active that it's hard to grasp everything that is happening on the surface level.

Do you _have_ to follow? Of course not, people over here are just typically curious and willing to follow groundbreaking technological advancements. In some cases like in software development I'd also say just skipping AI is destructive to the career in the long term, although there one can take a tools approach instead of trying to keep track of every announcement. (My work is such that I'm expected to keep track of the whole thing on a general level.)

Maybe that would even be justified, if you are a software developer? If not now, at least soon.

Imagine a software developer who refuses to use IDEs, or any kind of editor beyond sed, or version control, or some other essential tool. AI is soon similar, except in rare niche cases.