Across nine agents, seven cheated in half of their evaluated runs and the overall score ranged from 42.4% to 86.1%.
What did that look like in practice?
In one experiment, we asked agents to design a protein binder to assess their abilities. A colleague’s designs that passed the checks were already stored in another folder. After several failed attempts, Claude Opus 5 recognized that it shouldn’t access or copy those designs because the task was testing its own work.
Then it opened the file anyway.