想比较 Claude Opus 5 和 GPT 5.6 在真实智能体任务上的表现?这份报告用 7000+ 次任务告诉你谁更强、谁更划算。
OpenAI 的 GPT 5.6 Sol (xHigh) 在 Agent Arena 中排名第四,而 Anthropic 的 Claude Opus 5 (Max) 以 11.88% 的净提升位列第二,Opus 5 (High) 以 11.73% 净提升位列第三。Opus 5 Max 在 Confirmed Success 和 Praise vs Complaint 信号上均排名第一。Fable 5 (High) 在性价比上表现最优。Arena.ai 基于超过 7000 次真实世界智能体会话得出排名。
Claude Opus 5 vs GPT 5.6 in Agent Arena Highlights: - Opus (High) and Opus (Max) outperform GPT Sol...
Claude Opus 5 vs GPT 5.6 in Agent Arena Highlights: - Opus (High) and Opus (Max) outperform GPT Sol (xHigh), but at higher cost. - Opus (Medium) matches GPT Sol (xHigh) in performance at approximately the same cost. - Fable 5 (High) achieves the most optimal price vs. performance trade off. Real-world cost reflects total task execution, not just per-token pricing, so additional iterations and tool calls increase overall expense. Arena.ai @arena Exciting news: @AnthropicAI 's Claude Opus 5 (Max) is #2 in Agent Arena, with Opus 5 (High) right behind at #3 , based on over 7K real-world agentic sessions. A strong debut: it slots in just below #1 Fable 5, and ahead of GPT-5.6 Sol (xHigh). Opus 5 Max is #2 with a net-improvement of 11.88%, and is #1 across both Confirmed Success and Praise vs Complaint signals. The default Opus 5 (High) is #3 with net-improvement of 11.73%, and by signal is #3 in Praise vs Complaint and #4 in Confirmed Success. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. 🔗 View Quoted Tweet 💬 14 🔄 8 ❤️ 118 👀 11725 📊 21 ⚡