Anthropic发布Fable 5.1模型加入Arena评测

Fable 5.1 by @AnthropicAI is now in the Arena! Bring your toughest prompts and start voting. Scores...

精选理由

Anthropic新推Fable 5.1,能在智能体任务中使用多种工具,比平均模型表现更优。

AI 摘要

Anthropic的Fable 5.1模型现已加入Arena评测平台。该模型在数百万真实世界长周期智能体任务中进行测试。模型可访问网络搜索、文件系统和终端工具完成复杂工作流。评测使用因果追踪方法衡量模型相对于平均模型的表现。

原文 · lmarena.ai

Fable 5.1 by @AnthropicAI is now in the Arena! Bring your toughest prompts and start voting. Scores...

Fable 5.1 by @AnthropicAI is now in the Arena! Bring your toughest prompts and start voting. Scores coming soon. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. @claudeai Fable 5.1 is also available in: Code Arena: WebDev, Text, Vision, and Document. Claude @claudeai We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 6 🔄 8 ❤️ 150 👀 9726 📊 17 ⚡

Anthropic发布Fable 5.1模型加入Arena评测 · AI 热点