Register and share your invite link to earn from video plays and referrals.

Alex Shaw
@alexgshaw
Hacking on @terminalbench and @harborframework. Founding MTS @LaudeInstitute. Formerly Google. BYU alum.
Joined October 2021
771 Following    2.8K Followers
๐š‘๐šŠ๐š›๐š‹๐š˜๐š› ๐š›๐šž๐š— -๐š ๐š–๐šŽ๐š›๐šŒ๐š˜๐š›/๐šŠ๐š™๐šŽ๐šก-๐šŠ๐š๐šŽ๐š—๐š๐šœ-๐Ÿท-๐Ÿท
Today we're introducing APEX-Agents 1.1. As AI models advance, so do their methods to solve APEX-Agents tasks. Weโ€™re updating our benchmark with task specifications, tooling, and environments to maintain leaderboard accuracy. Most notably, APEX-Agents 1.1 no longer rewards noncommittal answers. This ensures that our tasks and grading align with how professionals operate. Claude Fable 5.1 remains on top of the leaderboard with 68.6% on Pass@1, though rankings have shifted throughout the rest of the board. Updated rankings for Pass@1: Claude Fable 5.1: 68.6% Gemini 3.7 Flash: 67.8% Claude Opus 5: 65.8% Grok 4.6: 65.3% GPT-6 Astra: 64.7% Read the announcement blog:
Show more