Agent Arena 发布智能体任务榜单:Anthropic 模型居 Code、Work、Chat 三类榜首
Arena.ai 更新了 Agent Arena 榜单,Anthropic 和 OpenAI 的模型在真实任务上谁强谁弱,一目了然。
Arena.ai 的 Agent Arena 基于 Agent 解决真实任务的在线记录,对模型在 Code、Chat、Work 三类智能体任务上排名。Anthropic 的 Claude Fable 5.1(Max)在 Code 和 Work 排第 1、Chat 排第 2,Claude Sonnet 5.5(Max)在 Chat 排第 1。OpenAI 的 GPT 6 Astra(Max)在三个类别均进前 4,Code 排第 2。Google 的 Gemini 4 Argon(High)Code 排第 14、Work 排第 12,但 Chat 排第 5,显示其在对话类任务上的差异化优势。榜单截至 2026 年 10 月 5 日。
The Agent Arena ranks agentic ability to solve problems across domains. U.S. labs lead across every domain: - @AnthropicAI models hold the #1 position in Code, Work, and Chat - @OpenAI 's GPT 6 Astra (Max) remains in the top four across all three categories: #2 in Code and #4 in both Work and Chat The leading models stay strong across categories, but their order shifts: - Claude Fable 5.1 (Max) leads both Code and Work and ranks #2 in Chat - Claude Sonnet 5.5 (Max) leads Chat and ranks #4 in Code and #5 in Work Code and Work rankings move together more closely than Chat: - Gemini 4 Argon (High), for example, ranks #14 in Code and #12 in Work but #5 in Chat, highlighting a distinct strength in conversational agent tasks Agent Arena ranks models based on overall net improvement against the current frontier, as well as performance across specific categories of agentic tasks. These rankings are grounded in live traces collected from agents attempting to solve real-world tasks submitted by humans around the world. - Code: writing and debugging code, workflow automation, and data analysis - Chat: creative writing, learning, personal questions, everyday research, and media generation - Work: documents, professional research, planning, and professional writing This visual outlines how selected models compare across Agent Arena categories as of October 5, 2026. Arena.ai @arena x.com/i/article/2089… 🔗 View Quoted Tweet 💬 12 🔄 1 ❤️ 62 👀 5882 📊 14 ⚡