想测智能体真实干活能力?Agent Arena用百万真实任务打分,比普通聊天评测靠谱多了。
Agent Arena发布新评测基准,在数百万个真实世界长程智能体任务上衡量模型表现。评测中模型需使用网络搜索、文件系统和终端工具完成复杂工作流,包括编写代码、制作幻灯片、网络研究、构建应用和文档分析。该基准采用因果追踪方法,可深入分析模型在长程任务中的决策过程。评测结果已上线arena.ai/leaderboard/ag榜单。
Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web se...
Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. Dive into the Agent Arena details at: arena.ai/leaderboard/ag… And learn more about our causal tracing methodology x.com/arena/status/2… Arena.ai @arena x.com/i/article/2085… 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 1 👀 756 📊 1 ⚡