想知道代码、工作、聊天分别谁最强?Agent Arena 新榜单给出答案:GPT 5.6 和 Claude Opus 5 各领风骚。
Agent Arena 新增代码、工作、聊天三个分类榜单。代码榜第一是 OpenAI 的 GPT 5.6 Sol(xHigh)。工作榜和聊天榜第一均为 Anthropic 的 Claude Opus 5(High 和 Max)。该评测基于数百万真实长时程智能体任务,单次长会话成本可达数百至一千美元。
New categories are now live in Agent Arena! The #1 ranked model is different in each category: - #...
New categories are now live in Agent Arena! The #1 ranked model is different in each category: - #1 for Code: GPT 5.6 Sol (xHigh) @OpenAI - #1 for Work: Claude Opus 5 (High) @AnthropicAI - #1 for Chat: Claude Opus 5 (Max) @AnthropicAI Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Across all categories, a single long agent session can cost hundreds or even a thousand dollars. Arena Researcher @evan_a_frick explains how: youtube.com/watch?v=dC7T8F… In this video, we're breaking down token pricing, caching, context growth, and why we use task-based cost estimates for leaderboards and Pareto graphs. 💬 1 🔄 0 ❤️ 5 👀 1851 📊 2 ⚡