Most AI benchmarks test what a model says. In commerce, the hard part was never the answer. It’s execution.
We’ve open-sourced CommerceAgentBench: a benchmark for real commerce operations.
Early results are humbling. The best overall completion rate is ~62%.
Qwen
@Alibaba_Qwen delivered the strongest overall performance across complex commercial workflows among the open-weight models evaluated.
Explore the benchmark and full results ↓