HN user

dontlikeyoueith

323 karma
Posts0
Comments155
View on HN
No posts found.

I disagree with your assessment pretty strongly -- the models themselves hit a wall over a year ago once companies exhausted all existing training data. LLMs don't induce world models, and they aren't capable of real search an planning outside their training distributions. They, structurally, never will be.

I haven't noticed a change in what I trust a model to generate in response to a single prompt in a year. The failure modes are unchanged. Yes, specific failures have improved as they have been documented and passed into model training data, but the way the models fail has not changed. They still fail for me nearly every single day. I'm a pretty heavy user - 3-4 Claude code processes running at a time, all day every day.

What has gotten better is tooling around the model -- but there's no space for exponential growth there. At least, not without exponential cost increase, which would make the whole thing untenable anyway.

Claude Fable 5 1 month ago

OpenAI and Anthropic are heavily subsidizing their inference -- no wait, they are charging the most they can get away with before going public. Where is the truth?

Both. They are charging the most they can get away with and that amount is still heavily subsidized by VC capital.

As an employer, I want AI to be fully allowed for assignments, and the assignments to be made trickier to compensate.

This is like saying first graders should learn to use calculators, not how to do arithmetic.

Some skills are foundational, and must be learned first in order to be able to solve harder problems. Skipping those skills because software can do them is moronic.

The NLA would be forced to use human readable representations to get a successful round trip.

That still doesn't guarantee any semantic correspondence between the human readable representation and the model's "thinking".

The child's game of "Opposite Day" is a trivial example of encoding internal thoughts in language in a way that does not correspond to the normal meaning of the language.

LLMs are based on neural networks, so one could create an interface where activating certain neurons triggers tool calls, with other neurons encoding the inputs; another set of neurons could be triggered by the tokenized result from the tool call.

You can do this. It's just sticking a different classifier head on top of the model.

Before foundation models it was a standard Deep RL approach. It probably still is within that space (I haven't kept up on the research).

You don't hear about it here because if you do that then every use case needs a custom classifier head which needs to be trained on data for that use case. It negates the "single model you can use for lots of things" benefit of LLMs.

GitHub Stacked PRs 3 months ago

I think the point the GP was trying to make is that the GitHub UI ought to be able to allow you to submit a branch with multiple well-organized commits and review each commit separately with its own PR

So the point he's trying to make is that Gituhub UI should support Stacked PRs but call them something else because he doesn't like the name?

GitHub Stacked PRs 3 months ago

Why do you insist on a different but functionally equivalent solution to the problem?

It's weird.

Why do we tolerate the fact that GitHub doesn't let you say "approved for changes in `frontend/*`

That's literally what stacked PRs are adding.

GitHub Stacked PRs 3 months ago

Depends on what you consider long-lived.

I typically generate stacks of 3-5 PRs in 1-2 days now (in a gen-AI world).

Nowhere is safe 3 months ago

There's been a massive step change in their capability per unit cost.

What used to cost millions per unit now costs tens of thousands. That's significant.

It's like saying artillery isn't that big a deal in 1914. After all, it's been around since 1452.

GPT-5.2 7 months ago

Refuse to express uncertainty or nuance (i asked ChatGPT to give me certainty %s which it did for a while but then just forgot...?)

They're literally incapable of this. Any number they give you is bullshit.