The ExploitGym paper evaluated several frontier models on the bench and reported that "Different models find different exploits" [1], so it seems most plausible that the "test solutions directly from Hugging Face’s production database" [2] which GPT-internal found were authored by Mythos (or some other LLM with complementary strengths), and placed in some internal HF repository when creating the ExploitGym paper/leaderboard.
[1] https://www.cybergym.io/exploitgym/#:~:text=Different%20mode...
[2] https://openai.com/index/hugging-face-model-evaluation-secur...