Agent Arena发布了各模型任务成本对比,能清楚看到不同模型花多少钱和效果,和同类比啥不一样。
Agent Arena对比了15款模型任务成本,Kimi K3 (Max)每任务成本仅$0.62;Claude Opus 5 (Max)成本达$3.37,与同类存在明显差距;Kimi K3 (Max)排名第4,成本远低于多数同级别模型;Grok和Qwen以较低成本实现了不错表现。
Top 15 models in Agent Arena vs. median task cost: who's delivering the most for their price? Key f...
Top 15 models in Agent Arena vs. median task cost: who's delivering the most for their price? Key findings: Even within the top 5, median cost per task ranges from $0.62 for Kimi K3 (Max) to $3.37 for Claude Opus 5 (Max), a more than 5x difference. Claude Opus 5 (Max) costs almost twice as much as Opus 5 (High), despite scoring slightly lower: - Opus 5 (Max): +12.0% at $3.37 per task - Opus 5 (High): +12.3% at $1.78 per task Kimi K3 (Max) stands out for value near the top. It ranks #4 with +10.5% net improvement, while its $0.62 median task cost is the lowest among the top eight. GPT-5.6 Sol (xHigh) is the highest-ranked OpenAI model at #5 . Its $1.39 median cost is lower than all three Anthropic models ranked above it, although its +9.8% score is also lower. Grok and Qwen deliver competitive performance at some of the lowest costs: - Grok 4.5 achieves +6.1% at $0.22 per task - Qwen-3.8 Max achieves +6.3% at $0.33 per task Agent Arena evaluates models on millions of real-world, long-horizon agentic tasks from a global community of users. Models use tools like web search, filesystem access, and terminal commands to complete complex workflows. Performance is measured as net improvement, using causal tracing methodology to estimate how much each model improves outcomes relative to the average model. Cost is measured below as median cost per task. 💬 9 🔄 3 ❤️ 48 👀 5293 📊 15 ⚡