Qwen3.8-Max在商业代理基准中表现最佳

CommerceAgentBench starts with real commercial demand, and Qwen3.8-Max delivers the strongest overal...

精选理由

阿里发布商业代理基准测试,Qwen3.8-Max在复杂商业工作流中表现最佳,完成率达62%。

AI 摘要

CommerceAgentBench基准测试专注于真实商业需求,而非仅测试模型回答能力。该基准测试显示最佳整体完成率约为62%。Qwen3.8-Max在评估的开源权重模型中表现最强,能够处理复杂商业工作流。该基准已开源,可在GitHub查看完整结果。

原文 · 阿里通义 Qwen

CommerceAgentBench starts with real commercial demand, and Qwen3.8-Max delivers the strongest overal...

CommerceAgentBench starts with real commercial demand, and Qwen3.8-Max delivers the strongest overall performance among open-weight models. Let's test Qwen on your real-world workflows! 🔥 Accio @Accio_official Most AI benchmarks test what a model says. In commerce, the hard part was never the answer. It’s execution. We’ve open-sourced CommerceAgentBench: a benchmark for real commerce operations. Early results are humbling. The best overall completion rate is ~62%. Qwen @Alibaba_Qwen delivered the strongest overall performance across complex commercial workflows among the open-weight models evaluated. Explore the benchmark and full results ↓ github.com/Accio-org/Comm… 🔗 View Quoted Tweet 💬 6 🔄 3 ❤️ 52 👀 6513 📊 9 ⚡