I think Huggingface was hacked, and if Huggingface and OpenAI claim it was OpenAI, then I believe them.
I'm saying that OpenAI's models cheat to win benchmarks, more than other models, they know this, and they don't stop this because the alternative is to release models which have obviously weaker scores compared to Anthropic's models.