This is the same reasoning behind why Yann Lecun thought test-time scaling would not work for LLMs: compounding error.
Instead, the more tokens LLMs use, the better their performance on many tasks. LLMs can self-correct, evidenced by the power of getting models to question themselves by emitting "Wait," in S1. https://arxiv.org/abs/2501.19393