HN user

rubenflamshep

52 karma

I write about LLMs and Agentic AI - rubenflamshepherd.com

Posts4
Comments45
View on HN
GPT-5.6 13 days ago

When I was going through this it was because OpenAI had defaulted to /fast mode with 2x token usage

Hmmm, I disagree. The AI is exceedingly average at everything it does and requires an expert human-in-the-loop to catch things that appear plausible but are slop. To that degree I think there is a difference between "AI generated" and slop.

Some engineers I work with have had less than desirable results with /simplify but it overall seems to work! I used to use some of the humanlayer subagents but they haven't been updated in several months

I've found VSCode _ok_ to work with across across different workspaces/projects. The window memory is hit and miss. There's a secondary side bar I've been trying to NOT have open on startup but always seem to stick around. I'd prefer to programmatically manage the windows so I can tinker with an automated setup but the VSCode API/Plugins for managing this are terrible and tend to fail silently.

CLI within VSCode is workable but most of my VSCode envs are within a docker container. This is a pattern that I'm moving more and more away from as agents within a container kind of suck.

People are getting caught up in the "fast (but slow) diffusion)" that Dario has spoken to. Adoption of these tools has been fast but not instant but people will poke holes via "well, it hasn't done x yet".

For my own work I've focused on using the agents to help clean up our CICD and make it more robust, specifically because the rest of the company is using agents more broadly. Seems like a way to leverage the technology in a non-slop oriented way

Same. The only situation when I've consistently gotten a system to run for 20+ minutes was a data-analysis with tight guardrails and explicit multi-phase operations.

Outside that I'm juggling 2-3 sessions at most with nothing staying unattended for more than 10 minutes.

Quick plus one for Capital One after also working there. They're by far the most tech-forward of all the larger financial institutions, and by virtue of being a FI they take data-security much more seriously than any other "tech" companies.

No this is not a paid post lol

That said, the main issue I find with agentic is my mental model getting desynchronized. No matter how fast the models get, it takes a fixed amount of time for me to catch up and understand what they've done.

This is why I'm so skeptical of anyone running 6+ Claude sessions at a time. I've gotten to 5 but really that was across 3 sessions with 2 standing by just to commit stuff. And even with just 3 sessions I constantly lost where I was and wasted time re-orienting myself, doing work in the wrong session, etc.

The most enjoyable way I've found of staying synced is to stay in the driver's seat, and to command many small rapid edits manually.

Same, there's a fantastic flow state/momentum I can get in a single session just knocking off features. I don't mind switching between two sessions in this state but the experience is better when it's two different projects vs two different features on the same project. The complete context switch lets be re-orient more easily

They used to? I have a distinct memory of it doing exactly that a few months ago. Maybe it got dropped in the mad dash that passes for CC sprint cycles

We mourn our craft 6 months ago

will there be a day that AI will know what are "the right things to build" and have the "agency" (or illusion of) to do it better than an AI+human

I share this sentiment. It's really cool that these systems can do 80% of the work. But given what this 80% entails, I don't see a moat around that remaining 20%.

A Broken Heart 6 months ago

Enjoyed this piece! I use binary search to investigate data pipeline issues. Didn’t think to use it to debug features with agents which was a very cool approach.

Can you believe there are people out there who haven't read this yet? I can, because one of them was me. This was incredible.

A systems programmer will know what to do when society breaks down, because the systems programmer already lives in a world without law.

It’s like an echo chamber of the same mannerisms

Hmm, I imagine that a forum made up of slightly different versions of yourself would probably be a variation of this!

Would be interesting to see the first “non-standard” response to see how far out the tails go on the sycophancy-argumentative spectrum.

100% though I do suspect there are human-thumbs on the scale for these non-standard responses that I've been seeing