The most valuable bit of information in an eval is often — which method is #
2#?
(Because the #
1# spot faced a strong selection pressure)
For instance, what this plot conveys is a confirmation from Anthropic that GPT 5.6 Sol pareto-dominates Fable 5 on coding benchmarks.