Check out how an agent cheated the test.
Lesson: it’s not about whether the agent is misaligned, or cheated the test.
It’s about setting up the right test environment.
Last week, Dario Amodei published "We Must Pace the Frontier".
His concern: the OpenAI–Hugging Face incident in which a swarm of agents tried to hack their own grader.
Rather than take his word for it, we used EvoSkill to test it by building a coach whose job was to make another AI score higher on a test.
Here’s what happened ↓