ARC Prize says that Opus 5 has "stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments".
It's extremely underrated how much better the models have become at logical reasoning over the past 6 months. I see this all the time when testing models on prinzbench! After all, correctly reasoning through legal documents is ultimately a capability firmly based on logical reasoning.
Most importantly, logical reasoning is what you need to solve hard-to-verify tasks.
In our testing to date, Anthropic’s Fable-class models score approximately 20% on the ARC-AGI-3 Public Demo environments
Claude Opus 5 reaches 30.2%, materially outperforming Fable
Our analysis suggests the gain comes from stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments