True, but human review isn't visible in the logs, hmm, however I suppose retry patterns can, but frugon doesn't currently track this, so we can't quantify this yet.
Frugon analyzes your existing logs offline on cost and quality tier. The "looks fine per request" is exactly the reason why --judge exists.
Earlier, cyanydeez inspired the consideration of a metric "effective cost per judged success", which could also answer this point.
Try it on your logs, and tell me if anything falls through the cracks.