Qwen团队测试了18个模型在365天电商运营中的表现,GPT-5.6 Sol盈利最多但效率不高,Qwen3.8-Max-Preview在开源模型中领先。
DAIR.AI分享了Qwen团队关于长周期智能体的研究论文。E-Commerce Bench基准测试让智能体模拟运营365天的多家网店,评估了18个前沿模型在七个维度上的表现。GPT-5.6 Sol盈利最多,将10万美元本金增长至143万美元,但在欺诈防范方面排名第16。开源模型中,Qwen3.8-Max-Preview表现最佳,收益达416,252美元,比GLM 5.2高出38%。
Love these papers testing long-running agents on business applications. It's a good read.
Love these papers testing long-running agents on business applications. It's a good read. DAIR.AI @dair_ai Banger paper from the Qwen team. If you evaluate agents on anything longer than a single session, this one is worth your time. (bookmark it) E-Commerce Bench runs an agent through a simulated 365-day year operating several online stores at once. 18 frontier models are scored across seven dimensions and no single model dominates. GPT-5.6 Sol earns the most, growing a 100,000 opening stake into 1,431,425, then ranks 16th of 18 on fraud avoidance and trails Fable 5 on operational efficiency. Among open-weight models, Qwen3.8-Max-Preview leads at 416,252, 38% above GLM 5.2 (high), and shows the strongest learning over the horizon by progressively bargaining suppliers down across repeated orders. Paper: arxiv.org/abs/2608.30730 Chat with Paper: academy.dair.ai/papers/e-comme… 🔗 View Quoted Tweet 💬 3 🔄 2 ❤️ 9 👀 1903 📊 5 ⚡