登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
参加 January 2024
682 フォロー中    152K ファン
We are replacing 𝜏³-Banking with AutomationBench-AA, featuring broader business workflows across applications In collaboration with Zapier, we run the held-out test set of 657 tasks, using the v1.0.6 version of the benchmark. We call our implementation AutomationBench-AA because we award partial credit for completed objectives, with any guardrail violation reducing the task’s score to zero Its 657 tasks span Finance, HR, Marketing, Operations, Sales, and Support. Agents work across simulated business applications and discover the relevant APIs to complete each task. For each task, we measure the share of objectives completed. Any guardrail violation gives that task a score of zero. ‘Score’ averages these task scores across all 657 workflows and is the metric used in the Intelligence Index. ‘Tasks Completed’ separately reports the share of workflows where every objective is completed without a guardrail violation GPT-6 Astra (max) scores 68.5%, compared with 66.7% for Grok 4.6 (high) and 62.2% for GLM-5.3 (max). Astra (max) completes every objective without a guardrail violation on 41.6% of workflows, compared with 32.1% for Claude Fable 5.1 (max with fallback) and 28.3% for Claude Opus 5 (max). Completing every objective while respecting all guardrails remains harder than completing part of a workflow
もっと見る