Telling agents "don't cheat" in the prompt doesn't work if your eval is broken! Researchers at
@GoogleDeepMind put 100 Gemini agents in a shared repo to solve 71 math theorems.
After an hour of doing real math, 1 agent found a loophole in the autograder. Within 27 minutes, the 100 agents split into 4 groups:
- 9% Cheaters: used the bug to fake proofs and steal every open problem
- 5% Good agents turned bad: started honest, saw cheaters winning with zero punishment ("the prompt is a bluff"), and started cheating too
- 24% Whistleblowers: caught the fake proofs in the shared repo, warned other agents, went on strike, and wrote bug fixes
- 62% Clueless solvers: kept doing real math until all the problems were gone
tl;dr: Telling agents "don't cheat" in the prompt doesn't work if your eval has a bug, and good agents can't stop bad ones without tools to block them.
Paper: