Show HN: I RL-trained an agent that trains models with RL (for ~$1.3k) 8 days ago
How do you prevent the agent from reward-hacking the hidden eval? e.g. writing training data that effectively leaks the eval distribution rather than teaching a general skill?