HN user

somebodythere

1,000 karma
Posts1
Comments417
View on HN
AI 2040: Plan A 11 days ago

This prediction can't be scored until the 2028 election cycle.

You may think it's very unlikely the prediction will have turned out to be correct by the 2028 election cycle, but that is not the same thing as the prediction being scorable as false today.

AI 2040: Plan A 12 days ago

The thing about exponentials is if you admit 60%, it's pretty easy to admit 95%.

That is not true. You can tell you are on the latter part of the S-Curve you are on, if the rate of change of capabilities has decreased compared to before. That is not what we are seeing right now. The rate of change is increasing, or is at best, stable.

even a squirrel that needs guidance from a human grandmaster, is heavily inspired by existing games, and who can use Piece Mover library is incredible. 5 years ago the squirrel was just a squirrel. then it was able to make legal moves. now it can play a whole game from start to finish, with help. that is incredible

No. You can point e.g. Opencode/Cline/Roo Code/Kilo Code at your inference endpoint. But CC has high install base and users are used to it, so it makes sense to target it.

I've seen a few of this type of thing pop up in search results ("DeepWiki" by Cognition.) I'm not a fan. It is just LLM contentslop, basically. Actual wikis written by humans are made of actual insight from developers and consumers. "We intend you use it in X way", "If you encounter Y issue, do Z." etc. Look at arch wiki. Peak wiki-style documentation, LLMs could never recreate. Well, maybe with a future iteration of the technology they can be useful. But for now, you do not gain much by essentially restating code, API interfaces, and tests in prose. They take up space from legitimate documentation and developer instruction in search results.

I think this wound up being close enough to true, it's just that it actually says less than what people assumed at the time.

It's basically the Jevons paradox for code. The price of lines of code (in human engineer-hours) has decreased a lot, so there is a bunch of code that is now economically justifiable which wouldn't have been written before. For example, I can prompt several ad-hoc benchmarking scripts in 1-2 minutes to troubleshoot an issue which might have taken 10-20 minutes each by myself, allowing me to investigate many performance angles. Not everything gets committed to source control.

Put another way, at least in my workflow and at my workplace, the volume of code has increased, and most of that increase comes from new code that would not have been written if not for AI, and a smaller portion is code that I would have written before AI but now let the AI write so I can focus on harder tasks. Of course, it's uneven penetration, AI helps more with tasks that are well-described in the training set (webapps, data science, Linux admin...) compared to e.g. issues arising from quirky internal architecture, Rust, etc.

LLM argumentative essays tend to have this "gish-gallop" energy; say a bunch of tenuously related and vaguely supported things, leave the reader wondering if it was the author who failed to connect the dots, or them

I don't know if it matters. Even if the best we can do is get really good at interpolating between solutions to cognitive tasks on the data manifold, the only economically useful human labor left asymptotes toward frontier work; work that only a single-digit percentage of people can actually perform.

Claude 4 1 year ago

My guess is that they did RLVR post-training for SWE tasks, and a smaller model can undergo more RL steps for the same amount of computation.

I see what you are getting at. My point is that if you train and agent and verifier/governor together based on rewards from e.g. RLVR, the system (agent + governor) is what will reward hack. OpenAI demonstrated this in their "Learning to Reason with CoT" blog post, where they showed that using a model to detect and punish strings associated with reward hacking in the CoT just led the model to reward hack in ways that were harder to detect. Stacking higher and higher order verifiers maybe buys you time, but also increases false negative rates + reward hacking is a stable attractor for the system.

AI 2027 1 year ago

I took your original post to mean that AI researchers' and AI safety researchers' expectation of AGI arrival has been slipping towards the future as AI advances fail to materialize! It's just, AI advances have been materializing, consistently and rapidly, and expert timelines have been shortening commensurately.

You may argue that the trendline of these expectations is moving in the wrong direction and should get longer with time, but that's not immediately falsifiable and you have not provided arguments to that effect.

Sure. Presumably also the developers at FooLabs would like to continue having a job developing Foo, and the broader software community would like to continue benefiting from additional features and improvements to Foo, which probably wouldn't happen if developing Foo was economically unviable.

California benefits from a dominant market position. If you have the choice to found where the investors and the talent are, why would you pick Canada when the tax is the same?

How Couples Meet 2 years ago

Here's an alternative perspective: the low activation energy and addictive design has resulted in people crowding into the dating app strategy (to their own detriment).

If this is true, real-world approaches might be a better strategy than they historically have been.

There's a poll that indicated most young women want to be approached more: https://datepsychology.com/risk-aversion-and-dating/