I built code-repair training data and shipped the eval so you can rerun it 4 days ago
you are Right let me rewrite it so it don't sound so try hard
Building TrueSET solo: code-repair training data where every example is proven by execution, not scored by another AI. I publish the numbers that make me look worse, too… building toward verified models delivered with the receipts to check them