Show HN: I had models play Slime Soccer against each other
https://slimeballbench.com/There were a few things that I wanted from this ai benchmark: * The models to compete directly with each other * To be able to _see_ how they were competing * Something more fun than a math test
Slime Soccer (a web game from my childhood) was like a good option since the decision set is small for each move and it felt nostalgic.
The most surprising thing is that the flagship models from each lab didn't always do best. For example, GPT 5.6 Sol was almost the worst performing model from OpenAI.