AI模型精选76°

Claude Sonnet 5 (Thinking) 在Agent Arena排行榜位列第6

Claude Sonnet 5 (Thinking) by @AnthropicAI debuts at #6 on the new Agent Arena leaderboard. In Agen...

精选理由

Anthropic 新出的 Claude Sonnet 5 主打智能体能力,Agent Arena 排第6,能自己规划并用工具完成任务。关心自主执行模型的可以看看它的实际表现。

AI 摘要

Anthropic 发布 Claude Sonnet 5 (Thinking),在 Agent Arena 排行榜上排名第6。该基准测试基于全球用户的数百万个真实长程智能体任务。模型可调用网络搜索、文件系统和终端等工具完成复杂工作流。Claude Sonnet 5 在任务成功率、用户满意度及 bash 能力上表现最强,工具幻觉率保持稳定。其可操控性分数置信区间较宽,仍在稳定中。

原文 · lmarena.ai

Claude Sonnet 5 (Thinking) by @AnthropicAI debuts at #6 on the new Agent Arena leaderboard. In Agen...

Claude Sonnet 5 (Thinking) by @AnthropicAI debuts at #6 on the new Agent Arena leaderboard. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. Claude Sonnet 5’s strongest signals are confirmed task success, praise vs. complaint, and bash capabilities. Tool hallucination holds stable. Note the wide confidence interval for steerability, as the scores are still stabilizing. See thread for more details on how Claude Sonnet 5 performs across 5 different signals. Claude @claudeai Introducing Claude Sonnet 5, our most agentic Sonnet yet. It makes plans, uses tools like browsers and terminals, and runs autonomously at a level that just a few months ago required larger and more expensive models. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 9 🔄 7 ❤️ 107 👀 12575 📊 16 ⚡

Claude Sonnet 5 (Thinking) 在Agent Arena排行榜位列第6 · AI 热点