we've barely scratched the surface when it comes to the available design space of agent harnesses.
i also expect you'll see markedly different results if you constrain yourself to small models. there even trivial harness improvements like Codex's /goal feature, and more capable basic tooling (e.g. semantic code grep, js-capable `fetch` tooling) make or break the actual task success rate.