We are replacing 𝜏³-Banking with AutomationBench-AA, featuring broader business workflows across applications
In collaboration with Zapier, we run the held-out test set of 657 tasks, using the v1.0.6 version of the benchmark. We call our implementation AutomationBench-AA because we award partial credit for completed objectives, with any guardrail violation reducing the task’s score to zero
Its 657 tasks span Finance, HR, Marketing, Operations, Sales, and Support. Agents work across simulated business applications and discover the relevant APIs to complete each task. For each task, we measure the share of objectives completed. Any guardrail violation gives that task a score of zero. ‘Score’ averages these task scores across all 657 workflows and is the metric used in the Intelligence Index. ‘Tasks Completed’ separately reports the share of workflows where every objective is completed without a guardrail violation
GPT-6 Astra (max) scores 68.5%, compared with 66.7% for Grok 4.6 (high) and 62.2% for GLM-5.3 (max). Astra (max) completes every objective without a guardrail violation on 41.6% of workflows, compared with 32.1% for Claude Fable 5.1 (max with fallback) and 28.3% for Claude Opus 5 (max). Completing every objective while respecting all guardrails remains harder than completing part of a workflow