Agent Arena发布了Pareto前沿平台,能查各模型做真实任务里的表现,像 Claude 和 GPT 这些对比后就知道哪个更划算。
Agent Arena推出的Pareto前沿平台上,Claude Opus 5(High)在真实代理任务中表现领先,提升+12.34%;该平台对比了各模型成本与性能,@OpenAI的GPT 5.5(High)表现突出,提升+7.75%;参与对比的多款知名模型包括Kimi K3(Max)和Groq 4.5等,帮助用户了解不同模型差异。
Pareto frontier is live in Agent Arena! Dive into model performance on real-world agentic tasks, com...
Pareto frontier is live in Agent Arena! Dive into model performance on real-world agentic tasks, compared to the median cost per task. The current models on the Pareto frontier for Agent Arena are: - Claude Opus 5 (High) by @AnthropicAI (+12.34%/ $1.78) - Kimi K3 (Max) by @Kimi_Moonshot (+10.53%/ $0.62) - GPT 5.5 (High) by @OpenAI (+7.75%/ $0.44) - GPT 5.5 by @OpenAI (+6.40%/ $0.28) - Grok 4.5 by @SpaceXAI (+6.08%/ $0.22) - GLM 5.2 (Max) by @Zai_org (+6.06%/ $0.18) - GPT 5.6 Luna (xHigh) by @OpenAI (+4.25%/ $0.04) - Mimo V2.5 Pro by @XiaomiMiMo (-2.22%/ $0.03) Check it out for yourself at the link below. Additional recent model releases will be landing on the Agent Arena leaderboard very soon. Stay tuned. Arena.ai @arena x.com/i/article/2089… 🔗 View Quoted Tweet 💬 7 🔄 1 ❤️ 19 👀 3236 📊 7 ⚡