Register and share your invite link to earn from video plays and referrals.

Chubby♨️
@kimmonismus
Tech Analyst & Content Creator. Editor-in-Chief @getsuperintel - Community of 300k+ in total: 🔗 //📧 kim@getsuperintel.com
Joined September 2022
3.4K Following    147.8K Followers
Most AI benchmarks test whether a model can give the right answer. CommerceAgentBench asks whether an agent can actually finish the work. Accio has open-sourced 107 e-commerce tasks across procurement, product listings, operations, fulfillment and after-sales. Agents work across browsers, email, calendars, documents, APIs and files. Crucially, they are not graded on what they claim to have done. The benchmark verifies what they actually changed, saved or submitted. Take the Gmail procurement case. The agent must search roughly 300 messy emails, identify the real suppliers, reconstruct the latest quotes, compare six Incoterms and four currencies, calculate landed costs and detect payment fraud. Then it must choose a supplier, label the relevant emails, save a reply draft and create a kickoff event. This is the kind of benchmark I find genuinely useful. It measures agents more like workers than chatbots. And the results show why human oversight still matters: the best observed run completed only 66 of 107 tasks, a 61.7% pass rate. And since 2026 is literally the year of agents, this is more important than ever. Accio says the tasks draw on 10M SMB users, 1.6M conversations, 200K agent trajectories and Alibaba’s 27 years of e-commerce experience. The project and task specifications are open source:
Show more