Alibaba Qwen amplified Accio’s open-source CommerceAgentBench and said Qwen3.8-Max posted the strongest overall result among open-weight models. The benchmark scores execution on real commerce operations rather than what a model says; early numbers put the best overall completion rate around 62%. Teams wiring agents into ordering, fulfillment, or support now have a public leaderboard for multi-step commercial work—and a reminder that even the leader still fails a large share of tasks.
Key Takeaways
- ✓CommerceAgentBench is open-source and scores commerce execution, not chat answers.
- ✓Best overall completion is about 62%; Qwen3.8-Max leads among open-weight models.
- ✓The gap is a signal that multi-step commercial agents still fail a large share of real workflows.