็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

Wade Foster
@wadefoster
Co-founder/CEO @Zapier
ๅ‚ๅŠ  January 2010
79 ใƒ•ใ‚ฉใƒญใƒผไธญ    25.9K ใƒ•ใ‚กใƒณ
Surprise drop today: Gemini 3.7 Flash. On AutomationBench, it beats models that cost twice as much. The progress here is insane. 3 weeks ago Gemini's 3.6 Flash scored 19.8%. Today, 3.7 is the first model to crack 30% on AutomationBench. At only 3 cents more per task. ๐—ช๐—ต๐—ฒ๐—ฟ๐—ฒ ๐—ถ๐˜ ๐˜„๐—ถ๐—ป๐˜€: Marketing (38%), Finance (37.5%), Sales (28.2%), and Support (22%) Example: Confirm a project is done in the CRM, then run our standard label cleanup on its email threads. Archive the closed ones, leave restricted ones alone. 3.7 was a full pass in 22 steps. GPT-5.6 Sol failed the same task in 8, then hallucinated a summary. ๐—ช๐—ต๐—ฒ๐—ฟ๐—ฒ ๐—ถ๐˜ ๐—น๐—ผ๐˜€๐—ฒ๐˜€: Operations. Opus 5 still runs that domain at 50%. Example: On a revenue-attribution task, 3.7 ran the math, updated the records, posted the summary, sent the escalation, and then never wrote the one required row on a second tracker. Came close, but still failed. How @Zapier's AutomationBench works: we score every new model on 657 of the hardest workflows we run. Scoring is deterministic: either the right records got updated and the right messages got sent, or they didn't (no partial credit) @GeminiAppโ€™s 30% is a new record. See every model and price here:
ใ‚‚ใฃใจ่ฆ‹ใ‚‹