Fable 5 重部署后各维度表现稳定,Agent 依旧最强,前端小幅波动属正常。想看前沿模型对比可以关注。
Arena.ai 收集了 Claude Fable 5 重新部署前后在 Text、Vision、Document、Code 和 Agent 五个 Arena 的数千投票,结果显示得分基本一致。Fable 5 在 Text、Document、Vision 和 Code Arena: Frontend 仍处前沿,Frontend 约 20 分下降在置信区间内。Agent Arena 中 Fable 5 曾排名第一,该评测基于数百万真实任务,使用因果追踪衡量模型表现。
The community has been asking how Claude Fable 5 compares before vs. after its latest re-deployment....
The community has been asking how Claude Fable 5 compares before vs. after its latest re-deployment. We collected thousands of votes on the new endpoint across Arenas - Text, Vision, Document, Code, and Agent - and here’s an early score preview. So far, scores look mostly consistent before and after re-deployment. Fable 5 remains at the frontier across Text, Document, Vision, and Code Arena: Frontend. The ~20-point drop in Frontend is still within the confidence interval as scores continue to stabilize. We’ll share more insights as more data comes in across all arenas - stay tuned! Arena.ai @arena Fable 5 is back in the Arena! When it first debuted, Fable 5 ranked #1 in Agent Arena: our benchmark for real-world, long-horizon agentic performance. Agent Arena evaluates models on millions of real tasks submitted by a global community of users, with access to web search, filesystem, and terminal tools to complete complex workflows. Rather than raw win rates, the leaderboard uses causal tracing to measure how much each model's outcomes outperform the average model. @AnthropicAI 's latest model is also live in Text, Vision, Document, and Code Arena: Frontend. Come give it a try in Battle Mode and Agent Mode and contribute to the leaderboards. 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 2 👀 1106 📊 1 ⚡