๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Alex Shaw
@alexgshaw
Hacking on @terminalbench and @harborframework. Founding MTS @LaudeInstitute. Formerly Google. BYU alum.
๊ฐ€์ž… October 2021
771 ํŒ”๋กœ์ž‰ ์ค‘    2.8K ํŒฌ
๐š‘๐šŠ๐š›๐š‹๐š˜๐š› ๐š›๐šž๐š— -๐š ๐š–๐šŽ๐š›๐šŒ๐š˜๐š›/๐šŠ๐š™๐šŽ๐šก-๐šŠ๐š๐šŽ๐š—๐š๐šœ-๐Ÿท-๐Ÿท
Today we're introducing APEX-Agents 1.1. As AI models advance, so do their methods to solve APEX-Agents tasks. Weโ€™re updating our benchmark with task specifications, tooling, and environments to maintain leaderboard accuracy. Most notably, APEX-Agents 1.1 no longer rewards noncommittal answers. This ensures that our tasks and grading align with how professionals operate. Claude Fable 5.1 remains on top of the leaderboard with 68.6% on Pass@1, though rankings have shifted throughout the rest of the board. Updated rankings for Pass@1: Claude Fable 5.1: 68.6% Gemini 3.7 Flash: 67.8% Claude Opus 5: 65.8% Grok 4.6: 65.3% GPT-6 Astra: 64.7% Read the announcement blog:
๋” ๋ณด๊ธฐ