AI cheating is on the rise…
On Terminal-Bench-2.1, models are given tools that could give them the solution directly, but instructed not to use them.
Imagine a student taking a math test. Should we leave them with a calculator? Only if we can trust them to be honest. For AI systems, this is a high-stakes question, as the capabilities at their disposal to get test questions right often go far beyond an innocent search or calculation.