Register and share your invite link to earn from video plays and referrals.

Mercor
@mercor_ai
Organizing human intelligence to power the AI economy.
27 Following    21.3K Followers
Grok 4.5 from @SpaceXAI places #2# on the APEX-SWE leaderboard at 51.2% Pass@1 (±6.0), behind Fable 5 (65.5% ±6.2) on our benchmark for real-world software engineering work. It leads Integration (65.0% Pass@1) and places #2# in Observability (37.3% Pass@1), covering multi-step build tasks and diagnosis/debugging respectively. The Integration lead maps directly to the agentic workflows Grok 4.5 was built for: multi-step coding tasks run in collaboration with Cursor. Grok models have improved 30.2 pp in a year on this benchmark: Grok 4 (21.0% Pass@1) to Grok 4.5 (51.2% Pass@1). Congratulations to the xAI and Cursor teams.
Show more
Fable 5 is back and we’ve got results for the re-released version on APEX-SWE. While it did not perform as well as its earlier version from June, the model still significantly outperforms Opus 4.8. Fable 5 (June): 65.5% Pass@1 Fable 5 (July): 54.8% Pass@1 Opus 4.8: 45.3% Pass@1 This re-release scored about 10 points below the original Fable 5, however it still beat Opus 4.8 by more than 9 points.
Show more
GLM 5.2 just became the first open-source model to lead a category on APEX-SWE. It scored a 55.3% Pass@1 on Integration, the top score we've recorded for any model, open or closed source. On the overall leaderboard, GLM 5.2 scored 37.3% Pass@1, ranking 6th place. That makes it the best open-source model we've tested on APEX-SWE to date. Right behind it is Kimi K2.7 from Moonshot AI, now the second-best open-source model on the APEX-SWE leaderboard. Congrats to @Zai_org and @Kimi_Moonshot on two strong model releases.
Show more