想知道你的模型在真实任务里啥水平?Agent Arena 用数百万实际任务和工具调用测试,结果很客观。
Agent Arena 的评分基于数百万来自全球用户的真实长期智能体任务。模型可使用网络搜索、文件系统和终端工具完成复杂工作流。排行榜使用因果追踪方法测量模型相对于平均模型的结果表现。详细方法可在 arena.ai/leaderboard/ag 查看。
How do we measure the performance in Agent Arena? The score is based on millions of real-world, lon...
How do we measure the performance in Agent Arena? The score is based on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. Head over to the Agent Arena leaderboard to see all the details: arena.ai/leaderboard/ag… 💬 0 🔄 0 ❤️ 6 👀 1497 📊 1 ⚡