HN user

yaodub

17 karma
Posts1
Comments9
View on HN
Claude Fable 5 1 month ago

SWE-Bench measures single tasks in isolation. In a real loop the model usually loses track of what I was trying to do long before code quality becomes the issue.

Claude Fable 5 1 month ago

Depends whether "unable to fully automate" means "needs occasional human checkpoints" or "slowly stops caring about your actual goal." Pretty different.