註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Mercor
@mercor
Organizing human intelligence to power the AI economy.
加入 April 2021
31 正在關注    24.5K 粉絲
GPT-6 Astra passes more tasks on APEX-Accounting than any other model. 13.1% Pass@1 (#1#) 60.0% mean score (#2#) Pass@1 is the proportion of tasks that a model scores 100% at least once across four attempts. Astra passes 56% more tasks than GPT-5.6 Sol and 12% more tasks than Fable 5.1. Most academic benchmarks measure model capabilities that are misaligned with real work. APEX benchmarks measure what enterprises actually care about. We built APEX-Accounting with @RampLabs to see if agents can handle a real company's books. Agents work the ledger in QuickBooks, tie it to bank statements in PDFs, chase figures across spreadsheets, and judge what is a real discrepancy. Results are graded against 2,186 criteria created by real accountants. See full leaderboard:
顯示更多
0
1
211
14
轉發到社區