Exciting news:
@AnthropicAI's Claude Opus 5 (Max) is #
2# in Agent Arena, with Opus 5 (High) right behind at #
3#, based on over 7K real-world agentic sessions. A strong debut: it slots in just below #
1# Fable 5, and ahead of GPT-5.6 Sol (xHigh).
Opus 5 Max is #
2# with a net-improvement of 11.88%, and is #
1# across both Confirmed Success and Praise vs Complaint signals. The default Opus 5 (High) is #
3# with net-improvement of 11.73%, and by signal is #
3# in Praise vs Complaint and #
4# in Confirmed Success.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.