Claude Opus 5.5 is the new #
1# on APEX-Agents and APEX-Accounting.
APEX-Agents:
73.5% Pass
@1 (#
1#)
81.3% mean score (#
1#)
APEX-Accounting:
15.4% Pass
@1 (#
1#)
62.0% mean score (#
1#)
On APEX-Agents, the new model gains +4.9 pp over Fable 5.1 (68.6%), the previous leader, and +7.6 pp over Opus 5 (65.8%). Anthropic says the biggest gains in this release are on long-running agentic tasks and knowledge work. That is what APEX-Agents measures, and the numbers agree.
Opus 5.5 on APEX domains:
Management Consulting: 80.0% Pass
@1 (#
1#)
Corporate Law: 71.2% Pass
@1 (#
3#)
Investment Banking: 69.3% Pass
@1 (#
2#)
Accounting: 62.0% mean score (#
1#)
Token use can indicate domains where the model is most effective. Consulting leads at 80% and uses the fewest tokens at 2.0M per attempt. Law has the highest partial credit (85.7% mean) but costs the most at 4.9M tokens. In Accounting, Opus 5.5 gets partial credit on most tasks but fully passes only 1 in 6.
More tokens can increase capability on Opus 5.5. Max effort consumes 3.50M tokens per attempt. That’s 2.1x more than Opus 5 and 1.2x Fable 5.1. But list price fell to $4/$20 per M (Opus 5 was $5/$25), and cache reads are $0.20, so 2.1x the tokens is only about 1.7x the dollars.
Effort level has a significant impact on benchmark scores and token usage.
Medium effort: 52.3% on 756k tokens
Max effort: 73.4% on 3.50M tokens
Increasing effort to max gains 21 points for 4.6x the tokens on agentic work. On APEX-Accounting, the same jump buys 2.8 points for 3.4x the tokens.
When Opus 5.5 solves a task, it solves it consistently. 142 of 239 tasks passed on all 4 runs.
Failing runs used 3.1M tokens vs 1.9M for passing runs, and took twice as long. 13 runs hit a hard failure, all in Corporate Law. 8 of them burned 25M to 42M tokens before dying, showing that long-horizon legal work can still send the model into a loop.
Congratulations to the
@claudeai team.
See full leaderboard: