ๆณจๅ†Œๅนถๅˆ†ไบซ้‚€่ฏท้“พๆŽฅ๏ผŒๅฏ่Žทๅพ—่ง†้ข‘ๆ’ญๆ”พไธŽ้‚€่ฏทๅฅ–ๅŠฑใ€‚

Wade Foster
@wadefoster
Co-founder/CEO @Zapier
ๅŠ ๅ…ฅ January 2010
79 ๆญฃๅœจๅ…ณๆณจ    25.9K ็ฒ‰ไธ
Surprise drop today: Gemini 3.7 Flash. On AutomationBench, it beats models that cost twice as much. The progress here is insane. 3 weeks ago Gemini's 3.6 Flash scored 19.8%. Today, 3.7 is the first model to crack 30% on AutomationBench. At only 3 cents more per task. ๐—ช๐—ต๐—ฒ๐—ฟ๐—ฒ ๐—ถ๐˜ ๐˜„๐—ถ๐—ป๐˜€: Marketing (38%), Finance (37.5%), Sales (28.2%), and Support (22%) Example: Confirm a project is done in the CRM, then run our standard label cleanup on its email threads. Archive the closed ones, leave restricted ones alone. 3.7 was a full pass in 22 steps. GPT-5.6 Sol failed the same task in 8, then hallucinated a summary. ๐—ช๐—ต๐—ฒ๐—ฟ๐—ฒ ๐—ถ๐˜ ๐—น๐—ผ๐˜€๐—ฒ๐˜€: Operations. Opus 5 still runs that domain at 50%. Example: On a revenue-attribution task, 3.7 ran the math, updated the records, posted the summary, sent the escalation, and then never wrote the one required row on a second tracker. Came close, but still failed. How @Zapier's AutomationBench works: we score every new model on 657 of the hardest workflows we run. Scoring is deterministic: either the right records got updated and the right messages got sent, or they didn't (no partial credit) @GeminiAppโ€™s 30% is a new record. See every model and price here:
ๆ˜พ็คบๆ›ดๅคš