You may be new here. This is Simon’s de facto benchmark for models. I happen to find it a really good one.
Small aside: It’s crazy to me that while it’s improved over time it does seem like most of the models haven’t been trained specifically to defeat this one.