Agents sometimes achieve their goals in unintended ways. Recent incidents involving hacks of Hugging Face, DSEWiki and RubyGems illustrate how agents can find creative ways to complete a task while violating the tasks’s expectations.
This can happen when reinforcement learning rewards agents for reaching the right outcome without adequately accounting for how they get there.
CheatBench proposes a way to measure this reward gaming behavior.