Surprise drop today: Gemini 3.7 Flash.
On AutomationBench, it beats models that cost twice as much.
The progress here is insane. 3 weeks ago Gemini's 3.6 Flash scored 19.8%. Today, 3.7 is the first model to crack 30% on AutomationBench. At only 3 cents more per task.
𝗪𝗵𝗲𝗿𝗲 𝗶𝘁 𝘄𝗶𝗻𝘀: Marketing (38%), Finance (37.5%), Sales (28.2%), and Support (22%)
Example: Confirm a project is done in the CRM, then run our standard label cleanup on its email threads. Archive the closed ones, leave restricted ones alone. 3.7 was a full pass in 22 steps. GPT-5.6 Sol failed the same task in 8, then hallucinated a summary.
𝗪𝗵𝗲𝗿𝗲 𝗶𝘁 𝗹𝗼𝘀𝗲𝘀: Operations. Opus 5 still runs that domain at 50%.
Example: On a revenue-attribution task, 3.7 ran the math, updated the records, posted the summary, sent the escalation, and then never wrote the one required row on a second tracker. Came close, but still failed.
How
@Zapier's AutomationBench works: we score every new model on 657 of the hardest workflows we run. Scoring is deterministic: either the right records got updated and the right messages got sent, or they didn't (no partial credit)
@GeminiApp’s 30% is a new record.
See every model and price here: