Claude Sonnet 5.5 (Max) 登上 Agent Arena 第三名,Anthropic 包揽前三
Anthropic 新版 Sonnet 5.5 排名很猛,Chat 类第一,编码只差 GPT-6 两分还便宜 20%,就是单任务贵了七成。
Claude Sonnet 5.5 (Max) 在 Agent Arena 首次上榜即排名第 3,净提升 +12.5%,比 Claude Sonnet 5 (High) 高出 8.1 个百分点。在 Chat 类别中它以 +15.6% 排名第 1,超过 Fable 5.1 的 +11.49% 和 Opus 5.5 的 +10.29%。代价是成本更高:单任务中位成本 $2.74,比排名第 2 的 Claude Opus 5.5 (High) 的 $1.58 高约 73%。在 Code Arena 的 WebDev 方向,xHigh 推理版以 1786 分排第 3,距第 2 名 GPT-6 Astra 的 1788 分只差 2 分,价格为其 80%。目前 Anthropic 包揽 Agent Arena 前三名。
Exciting news: Claude Sonnet 5.5 (Max) by @AnthropicAI has debuted at #3 in the Agent Arena with +12.5% net improvement! This release is a 8.1 percentage-point increase over Claude Sonnet 5 (High), which ranks #13 with +4.4% net improvement. By category, Claude Sonnet 5.5 secured the #1 spot in Chat (+15.6%) above both Fable 5.1 (+11.49%) and Opus 5.5 (+10.29%). This performance comes with a higher cost: Claude Sonnet 5.5 (Max) has a median cost of $2.74 per task, about 73% higher than #2 Claude Opus 5.5 (High) at $1.58. @AnthropicAI models now hold all three top positions in Agent Arena. Congrats to the team! Arena.ai @arena Exciting update: Claude Sonnet 5.5 with xHigh reasoning has landed in the Code Arena: WebDev. With 1786 pts, its ranked #3 ! At a blended $8/M tokens, Claude Sonnet 5.5 remains on the Pareto frontier with xHigh reasoning. This release is just 2 pts from GPT-6 Astra in the #2 spot with 1788 pts, for 80% of the price. By domain, Claude Sonnet 5.5 (xHigh) landed: - #2 in Gaming, Reference-Based Design, and Brand & Marketing - #3 in Simulations - #4 in Content Creation Tools and Consumer Product - #6 in Data & Analytics Congrats again to @AnthropicAI on this release! 🔗 View Quoted Tweet 💬 25 🔄 22 ❤️ 360 👀 44162 📊 50 ⚡