Agent Arena 发布新基准测试,评估模型在百万真实长周期任务中的表现
Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web se...
Agent Arena 发布了新基准测试,用百万真实长周期任务评估模型,比之前的基准更贴近实际应用场景。
Agent Arena 测量模型在百万真实长周期代理任务上的表现。模型获得网页搜索、文件系统、终端等工具,完成复杂工作流,如写代码、创建演示文稿、研究网页、构建应用和分析文档。使用因果追踪方法测量模型的净改进,反映其相对于平均模型的成果提升程度。
Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web se...
Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Dive into Agent Arena at: arena.ai/leaderboard/ag… and check out the Pareto frontier: arena.ai/leaderboard/ag… 💬 1 🔄 0 ❤️ 2 👀 1273 📊 1 ⚡