Register and share your invite link to earn from video plays and referrals.

prinz
@deredleritt3r
ad astra | | prinzbench:
Joined January 2024
4.5K Following    21.4K Followers
ARC Prize says that Opus 5 has "stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments". It's extremely underrated how much better the models have become at logical reasoning over the past 6 months. I see this all the time when testing models on prinzbench! After all, correctly reasoning through legal documents is ultimately a capability firmly based on logical reasoning. Most importantly, logical reasoning is what you need to solve hard-to-verify tasks.
Show more
In our testing to date, Anthropic’s Fable-class models score approximately 20% on the ARC-AGI-3 Public Demo environments Claude Opus 5 reaches 30.2%, materially outperforming Fable Our analysis suggests the gain comes from stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments
Show more