HN user

fourthark

1,268 karma

I code

Posts16
Comments582
View on HN

Interesting use of evals.

Might help interpretation to say on the front page that it's a five point scale with 0 (or 1?) being the safest score. This can be picked up from colors and the bars in the individual reports, but it takes a minute to figure it out.

Interesting, I got a completely different result on green/blue on this one, way more green whereas I got average on the individual test. Going between very different colors makes it hard to reset - they might consider breaks between spectra.

You gasp. You hyperventilate. Your heart rate jumps. Your blood pressure climbs. All of this in a few seconds.

There's something especially creepy about AIs talking in the second person about biological processes they don't experience.

Arguing with Agents 3 months ago

reset the context

Yes. Do this. These problems likely mean you have muddled the context.

The article too long and I didn't read the whole thing, but I'm glad the author came to understand that arguing won't help.

I think the point is that you have a better idea of what you want it to remember and even a small hint can have big impact.

Just saying "write up what you know", with no other clues, should not perform any better than generic compaction.

Odd how this thread is a recapitulation of your experience with the LLM.

What is take from this is that it's pointless to try to find out why an LLM does something - it has no intentions. No life and no meaning, quite literally.

And if you try to dig you'll only activate other parts of its training, transcripts of people being interrogated - patients or prisoners, who knows. Scary and uncreative stuff.

Strongly agree about the deterministic part. Even more important than a good design, the plan must not show any doubt, whether it's in the form of open questions or weasel words. 95% of the time those vague words mean I didn't think something through, and it will do something hideous in order to make the plan work

Why do you find it useless for legacy code? I find I have to give it plenty of context but it does pretty well on legacy code.

And Ask DeepWiki is a great shortcut for finding the right context… Granted this is open source and DW is free.

Is it the specific nature of your work?

But only if there is a competent compiler engineer running the AI, reviewing specs, and providing decent design goals.

Yes it will be far easier than if they did it without AI, but should we really call it “produced by AI” at that point?