가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Mercor
@mercor
Organizing human intelligence to power the AI economy.
가입 April 2021
31 팔로잉 중    24.5K 팬
GPT-6 Astra passes more tasks on APEX-Accounting than any other model. 13.1% Pass@1 (#1#) 60.0% mean score (#2#) Pass@1 is the proportion of tasks that a model scores 100% at least once across four attempts. Astra passes 56% more tasks than GPT-5.6 Sol and 12% more tasks than Fable 5.1. Most academic benchmarks measure model capabilities that are misaligned with real work. APEX benchmarks measure what enterprises actually care about. We built APEX-Accounting with @RampLabs to see if agents can handle a real company's books. Agents work the ledger in QuickBooks, tie it to bank statements in PDFs, chase figures across spreadsheets, and judge what is a real discrepancy. Results are graded against 2,186 criteria created by real accountants. See full leaderboard:
더 보기