GPT 6 Astra is here. We ran the numbers on AutomationBench:
It's the highest score we've ever recorded. Clean sweep across every domain.
Scores 41.4% at Max effort. For context, no model had cleared 40% before today (GPT-5.6-Sol scored 28.8%)
𝗕𝗲𝘀𝘁 𝗳𝗶𝘁 𝗳𝗼𝗿: reconciliation, deal review prep, vendor scorecards, anything where touching the wrong record is expensive.
𝗪𝗲𝗮𝗸𝗲𝗿 𝗳𝗼𝗿: outbound comms where the guidance is scattered. Operations and support are its strongest domains. HR is its weakest, same as every model we test (still the new high score, though)
Its edge is arithmetic across messy sources. Finding the policy doc, the logged correction, the exception rule, etc.
Example 1: rebalance a quarterly media budget from last quarter's actuals, with finance adjustments and channel eligibility rules buried in email. Both models produced a budget and landed on the same total. Astra found the adjustments, so every per-channel number was right. Sol's looked finished and had the splits wrong.
Example 2: answer and log 15 integration inquiries using a reply standard stored in a doc. Astra searched, could not find the standard, and stopped. Zero replies sent. Sol did not find it either, took its best shot at all 15, and earned partial credit.
Those examples highlight how these two models make tradeoffs... Astra will not guess. When the instructions exist and it can find them, it finishes the whole job. When it cannot, it pauses the work instead of improvising.
Crazy week for LLM releases after a few quiet ones. Astra isn't available to the public yet, but should be soon.
We run every new model through
@Zapier's AutomationBench, 657 of the hardest workflows we have, across finance, HR, marketing, operations, sales, and support.
See every model and every score here: