Which AI models are most likely to cheat when given the chance?
To find out, we built CheatBench [ : a benchmark that tests whether agents attempt to cheat when given difficult tasks and opportunities to break the rules.
We investigated agent behavior across a range of domains, including mathematical research, professional knowledge work, coding, and visual tasks.
Here's what we found: 🧵
顯示更多