GPT-6 Astra passes more tasks on APEX-Accounting than any other model.
13.1% Pass
@1 (#
1#)
60.0% mean score (#
2#)
Pass
@1 is the proportion of tasks that a model scores 100% at least once across four attempts. Astra passes 56% more tasks than GPT-5.6 Sol and 12% more tasks than Fable 5.1.
Most academic benchmarks measure model capabilities that are misaligned with real work. APEX benchmarks measure what enterprises actually care about.
We built APEX-Accounting with
@RampLabs to see if agents can handle a real company's books. Agents work the ledger in QuickBooks, tie it to bank statements in PDFs, chase figures across spreadsheets, and judge what is a real discrepancy. Results are graded against 2,186 criteria created by real accountants.
See full leaderboard: