가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Mercor
@mercor_ai
Organizing human intelligence to power the AI economy.
가입 April 2021
27 팔로잉 중    21.3K
Grok 4.5 from @SpaceXAI places #2# on the APEX-SWE leaderboard at 51.2% Pass@1 (±6.0), behind Fable 5 (65.5% ±6.2) on our benchmark for real-world software engineering work. It leads Integration (65.0% Pass@1) and places #2# in Observability (37.3% Pass@1), covering multi-step build tasks and diagnosis/debugging respectively. The Integration lead maps directly to the agentic workflows Grok 4.5 was built for: multi-step coding tasks run in collaboration with Cursor. Grok models have improved 30.2 pp in a year on this benchmark: Grok 4 (21.0% Pass@1) to Grok 4.5 (51.2% Pass@1). Congratulations to the xAI and Cursor teams.
더 보기