Agent Arena 以数百万真实长程任务评测智能体模型

Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web se...

精选理由

想知道哪个模型最能搞定真实长任务?Agent Arena 用几百万个任务跑分,还给模型上网、文件、终端工具,比谁净提升大。

AI 摘要

Agent Arena 使用数百万个真实世界长程智能体任务评测模型。模型获得网页搜索、文件系统和终端工具,完成写代码、制作幻灯片、网络研究、构建应用和分析文档等工作流。评测采用因果追踪方法衡量模型的净改进,即相对平均模型提升结果的程度。完整排行榜见 arena.ai/leaderboard/ag…。

图片来源 · lmarena.ai
原文 · lmarena.ai

Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web se...

Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Learn more at: arena.ai/blog/agent-are… See the full Agent Arena leaderboard: arena.ai/leaderboard/ag… 💬 0 🔄 0 ❤️ 2 👀 304 ⚡