Seems like I have much better results with Gemini <= 3.5 Flash being able to solve 10 Lake levels. Of course there may be methodology differences (like how many times the level is attempted from scratch).
I wonder if the reasoning tokens are still being accidentally discarded between turns in those tests. It seems to be much more important to preserve those for Gemini, as it likes to stay quiet and and just make tool calls, unlike Claude that yaps a lot what it's planning to do.