Agent Arena发布了新模型评测标准,用多种工具测模型实际能力,比普通评测更全面。
Agent Arena推出新模型评测标准,基于百万级真实世界长期任务;利用因果追踪方法测试模型在网页搜索等工具的复杂流程;通过与平均模型对比评估任务效果。
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a globa...
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. See the full Agent Arena leaderboard: arena.ai/leaderboard/ag… 💬 0 🔄 1 ❤️ 2 👀 1269 📊 1 ⚡