看看 arena.ai 上新出的 Agent Arena 排行榜,它用真实任务测试模型,还能用工具,比普通测试更实用。
Agent Arena 排行榜衡量模型在数百万个真实世界、长周期任务中的表现。模型可使用网络搜索、文件系统和终端工具完成复杂工作流。该榜单通过因果追踪方法,衡量模型相对于平均模型的结果表现。
Check out the Agent Arena leaderboard to see the details: https://t.co/qNgm4b4XHl In Agent Arena, w...
Check out the Agent Arena leaderboard to see the details: arena.ai/leaderboard/ag… In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. 💬 0 🔄 0 ❤️ 0 👀 825 ⚡