行业83°

Claude Fable 5 因美国政府指令从 Arena 下架,此前排名第一

Update: We've removed Claude Fable 5 from Arena, following Anthropic's latest announcement and the U...

精选理由

最强模型被下架,原因值得关注

AI 摘要

Arena 宣布已移除 Claude Fable 5,原因是 Anthropic 的最新公告和美国政府指令要求暂停访问。Fable 5 在 Agent、Text 和 Code Arena 三项基准中均排名第一,是 Arena 测试过的最强模型,在 Agent Arena 上以最大领先幅度超过 Opus-4.8 和 GPT-5.5。该模型在确认任务成功率和好评/投诉比两项关键信号上表现突出,但可操控性较弱。Arena 表示将在可能时恢复访问并重启社区测试。

原文 · lmarena.ai

Update: We've removed Claude Fable 5 from Arena, following Anthropic's latest announcement and the U...

Update: We've removed Claude Fable 5 from Arena, following Anthropic's latest announcement and the U.S. government directive to suspend access. Claude Fable 5 is the most powerful model we’ve ever tested - ranking #1 across Agent, Text, and Code Arena, and setting a new breakthrough for frontier AI performance. We look forward to restoring access and resuming community testing when possible. Arena.ai @arena Exciting news: Claude Fable 5 ranks #1 on the new Agent Arena leaderboard! Fable 5 leads by the widest margin ever over Opus-4.8 and GPT-5.5 on two key signals: confirmed task success rate and praise vs. complaint, despite weaker steerability. If Fable can do something, it will do it very well. If it can't/doesn't want to do something, it may be hard to steer the model towards the goal. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide deck, researching the web, building apps, and analyzing documents. We use the causal tracing methodology to measure a model's net improvement which indicates how much it improves outcomes relative to the average model. Huge congrats to @AnthropicAI for the incredible milestone! Below we break down how Claude Fable 5 (based on Mythos) scored across 5 signals, drawn from tasks submitted by a global community of users. 🔗 View Quoted Tweet 💬 11 🔄 18 ❤️ 215 👀 23302 📊 29 ⚡