模型多源确认78°

Claude Fable 5.1 (Max)登顶Agent Arena榜首

Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1! At a $4.14 median cost p...

精选理由

Anthropic的Claude Fable 5.1 (Max)在Agent Arena排名第一,性能领先但成本较高,用户满意度极高。

Claude Fable 5.1 (Max)在Agent Arena排行榜中位列第一,净性能提升达15.8%。该模型基于6.7万+真实智能体会话测试,每任务中位成本为4.14美元。用户情感反馈显示,其表扬率比投诉率高42.5%,是第二名模型的2倍。

原文 · lmarena.ai

Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1! At a $4.14 median cost p...

Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1 ! At a $4.14 median cost per task and a +15.8% net improvement, it has reshaped the Pareto frontier! Claude Fable 5.1 is both the most performant and costly among all Agent Arena models today. Based on 6.7k+ real-world agentic sessions, Claude Fable 5.1 stands out in implicit user sentiment with an astonishing (+42.5%) in Praise vs. Complaint. Users are praising it around 2x more often than the next top model. It also sees strong explicit feedback via Confirmed Success (+22.4%), and solid Bash Recovery (+13.1%), with no Tool Hallucinations. Stay tuned as traces continue to come in for the latest GPT-6 Astra to see how it compares. Use Agent Mode to contribute to the real-world rankings. Arena.ai @arena Claude Fable 5.1 (Max) by @AnthropicAI has landed in the Agent Arena at #1 with +15.8% net improvement across 6.7k+ real-world agentic sessions! It also redraws the price-performance frontier: #1 on the leaderboard at a median cost of $4.14/task. By signal, Claude Fable 5.1 sees a massive lead in implicit user sentiment with an astonishing (+42.5%) in Praise vs. Complaint. Users are praising it around 2x more often than the next top model. It also sees strong explicit feedback via Confirmed Success (+22.4%), and solid Bash Recovery (+13.1%), with no Tool Hallucinations. More detail on its performance by signal below. Claude Fable 5.1 (Max) out ranks all past Claude variants and the rest of the pack by a healthy lead. In Agent Arena, we measure models on millions of long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. Stay tuned as traces continue to come in for the latest GPT-6 Astra to see how it compares. Use Agent Mode to contribute to the real-world rankings. Congrats again to @AnthropicAI for this release. 🔗 View Quoted Tweet 💬 4 🔄 0 ❤️ 11 👀 2809 📊 4 ⚡