HN user

mrbungie

1,423 karma
Posts0
Comments542
View on HN
No posts found.

Just add a --verbose flag that shows the stacktrace when there is an error. Then add a footer message when an error appears in non-verbose mode that invites the user/agent to use --verbose to get the full picture.

It obviously may end up in thousands of tokens burned through though (you can also fix that adding different levels of verbosity), but hopefully errors are not common.

Claude Sonnet 5 22 days ago

From what I gather from GPs upper post: Technical debt, skill atrophy, delusions of grandeur about one's own abilities / psychosis.

You don't need SOTA-level LLMs to create value with AI. Hell, you can build good solutions with a simple small finetuned models.

When models are good, expectations are adjusted accordingly to deliver things on par with the whole industry, you can't just say, I have built my own Intel Pentium II, now I will try to use it to compile Electron App and run 3DS Max there.

I know you are taking your analogy to its breaking point but it really depends on what you are doing. I know people that use 10+-year old thinkpads and they do just fine.

That's convenient accounting. The reality is that they can't stop training since they risk losing customers if they do so. So they shouldn't factor it out of profitability analysis.

AI as a tech is fine. But disliking it and the social/economic effects around it is fine too, people should be allowed to feel however they want to feel about certain techs and situations.

To recommend people to suck it up is not the answer I wish in the society I want to live in.

Gemini 3.5 Flash 2 months ago

I know artificial analysis quite well as the gold standard in llm evals.

I also know them, but it took me a while to realise you were publishing their data in that table. I don't think it was clear.

The age is important because new techniques keep being developed and so it is a very rough indicator of the size/cost/efficiency trade-off.

Yes but you are already including the name of the model, your potential public for the table already know about model's release history and therefore each model's age, at least roughly.

Gemini 3.5 Flash 2 months ago

It is kind of noisy because the release recency, which is what your "age" column actually represents, is not important data for the comparison you are trying to make.

Also what message we should get from that table is not really obvious.

Well, we were overall better for a couple of centuries after abolishing all-powerful kings + some welfare laws here and there (ymmv, maybe serfdom sounds nice to you). So those changes can work for a while, big emphasis on can and for a while.

Greedy accumulators always end up ruining things for societies when it gets into ridiculous extremes (and there is a part of society that notices and gets fed up).

I still don't know how to reconcile these reports with what other people say about GenAI-agentic assisted engineering being the only way of working nowadays, especially in startups.

Probably there is no dichotomy going on and it depends on multiple factors, but it seems so weird to see reports that are so different between each other.

I'm always wondering who has the time to consume all the new code that is being produced. Like sure, you can produce at 5-10X the speed, but is someone using those features? Not sure if the typical consumer mind can keep up with such speed of changes.

You usually don't get immediate responses from hires which means delayed gratification and avoiding much of the potential dopaminergic effects you get when engaging with LLMs.

You can play overextending the hire analogy all you want but it is simply not the same.

The gambling part is because of the (hopefully emergent and not purposefully designed) intermittent reinforcement due to the limits. You don't get that with regular hires.

Is psychohistory a thing now? They are just another actor, for sure they have relevant usage data but (1) their view are biased in multiple ways (at least due to perverse incentives), knowingly and unknowingly (2) they are experts in their field, but not necessarily in sociology, economy, and politics, which are arguably necessary to make serious, even if ineffective, predictions about what's going to happen with labor.

I'd rather see what institutions such as NBER have to say than to blindly trust frontier AI labs that won't even publish their methodologies most of the time.