가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

prinz
@deredleritt3r
ad astra | | prinzbench:
가입 January 2024
4.5K 팔로잉 중    21.4K
ARC Prize says that Opus 5 has "stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments". It's extremely underrated how much better the models have become at logical reasoning over the past 6 months. I see this all the time when testing models on prinzbench! After all, correctly reasoning through legal documents is ultimately a capability firmly based on logical reasoning. Most importantly, logical reasoning is what you need to solve hard-to-verify tasks.
더 보기
In our testing to date, Anthropic’s Fable-class models score approximately 20% on the ARC-AGI-3 Public Demo environments Claude Opus 5 reaches 30.2%, materially outperforming Fable Our analysis suggests the gain comes from stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments
더 보기