Aggregating across release dates shows a steady increase in cheating across coding benchmarks, particularly among newer model families.
As labs compete to produce more powerful models, RL environments are not always carefully audited. When environments allow cheating, models can be reinforced on the correctness of this behavior. Worse still, model providers may use the same infrastructure to prevent cheating in both training and evaluation—so if models learn to evade those safeguards, that behavior may transfer and inflate evaluation results.