Dario's essay says "We must slow the pace at which we improve the capabilities of AI models." So what did Anthropic commit to?
If you believed that sentence, you would expect something that touches compute: training runs capped at half or a quarter of the lab's compute, compute moved from training to inference, or a ban on a form of development, like using AI to do AI research. On recursive self-improvement the essay says only that it "must be pursued very carefully, if at all."
Instead, Anthropic committed to third-party evaluators with employee-level access (desks, badges, laptops) and the right to publish what they find. Everything else in the essay is a request that other people act: other labs, governments, export controls.
Evaluators watch what a lab does. They do not change how much compute it buys. So on its face this is regulatory capture done well. You position yourself as the most aligned lab, your rivals fall in line within hours, and the regulation gets written to your spec. It looks like Dario played 4D chess and won.
However, I don't think he did, and I explain why below.