Anthropic 的 Opus 5 在 Agent Arena 上超过 GPT-5.6,你也能用自己提示词在 agent 模式里测试它。
Anthropic 的 Claude Opus 5 (Max) 在 Agent Arena 中以 11.88% 净改进率位居第二,紧随其后的 Opus 5 (High) 以 11.73% 位列第三。该排名基于超过 7,000 次真实世界智能体会话,且 Opus 5 Max 在 Confirmed Success 和 Praise vs Complaint 信号上均获得第一。默认版本 Opus 5 (High) 在 Praise vs Complaint 信号排第三,在 Confirmed Success 信号排第四。Agent Arena 通过数百万个长期任务评估模型的工具调用和复杂工作流完成能力。
More from Opus 5 and @petergostev. Put Opus 5 to the test with your own prompts in Agent Mode!
More from Opus 5 and @petergostev . Put Opus 5 to the test with your own prompts in Agent Mode! Your browser does not support the video tag. 🔗 View on Twitter Arena.ai @arena Exciting news: @AnthropicAI 's Claude Opus 5 (Max) is #2 in Agent Arena, with Opus 5 (High) right behind at #3 , based on over 7K real-world agentic sessions. A strong debut: it slots in just below #1 Fable 5, and ahead of GPT-5.6 Sol (xHigh). Opus 5 Max is #2 with a net-improvement of 11.88%, and is #1 across both Confirmed Success and Praise vs Complaint signals. The default Opus 5 (High) is #3 with net-improvement of 11.73%, and by signal is #3 in Praise vs Complaint and #4 in Confirmed Success. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. 🔗 View Quoted Tweet 💬 1 🔄 3 ❤️ 10 👀 1961 📊 2 ⚡