How do you catch AI agent regressions after prompt or model changes?

https://news.ycombinator.com/item?id=48313367
by 1taimoorkhan0 • 2 months ago
2 1 2 months ago

Seeing a pattern where teams fix a failure in an agent, change the prompt or model a week later, and the same failure quietly comes back. Nobody catches it until a user does. Curious how people are handling this today. Manual test cases? Evals? Logs? Nothing? Not trying to pitch anything. Just trying to understand how widespread this is and what current approaches look like.

Related Stories

Loading related stories...

Source preview

news.ycombinator.com