想看看人机协作到底有没有用?Agent Arena拿数据说话,GLM-5.2开源最强,Claude Fable 5刚登顶就被叫停,这瓜值得吃。
Agent Arena推出了因果追踪方法论,通过分析人类与AI代理协作的追踪数据来量化协作的真实价值,并能观测到广泛的模型行为。基于该方法的新排行榜显示,GLM-5.2 (Max)进入前十,成为最强开源模型,确认成功率比基线高+9.4%,表扬-抱怨比高+14.9%。Claude Fable 5在几乎所有指标上曾排名第一,但因美国政府指令暂停访问。排行榜基于数百万个真实世界长期代理任务,使用因果追踪评估模型相对于平均模型的表现。
Agent Arena's causal tracing methodology lets us quantify the real value of humans working together ...
Agent Arena's causal tracing methodology lets us quantify the real value of humans working together with AI agents, and observe a huge range of model behaviors from the same traces. We started with 5 signals: confirmed success, praise vs. complaint, steerability, bash recovery, and tool hallucination. But the surface area is effectively unbounded, there's so much more to explore. Stay tuned. Listen in as @ml_angelopoulos and Evan get into what's possible. 👇 Your browser does not support the video tag. 🔗 View on Twitter Arena.ai @arena Agent Arena has been live for 2 weeks, with 10 more models now on the new leaderboard. Two highlights worth mentioning: - GLM-5.2 (Max) by @Zai_Org enters the top 10. The strongest open-weight result we've measured, at +9.4% confirmed success and +14.9% praise-vs-complaint relative to baseline. - Claude Fable 5 by @AnthropicAI debuted at #1 across nearly every metric before the U.S. government directive to suspend access. It’s a useful upper bound for where the frontier currently sits. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. Which model will enter the Arena next? Read more about the methodology and check out the live leaderboard (links in thread) 👇 🔗 View Quoted Tweet 💬 4 🔄 4 ❤️ 14 👀 2279 📊 5 ⚡