AI模型精选

Agent Arena发布token效率分析:Opus与Fable表现突出

[Token efficiency in Agent Arena] Agent Arena measures agent performance across a range of real-worl...

精选理由

想找token性价比高的模型?Agent Arena告诉你Opus和Fable有多能打,GPT-5.5也很省token。

AI 摘要

Agent Arena通过代码编写、幻灯片制作等真实任务评估模型性能。Opus 4.8 Thinking每会话消耗较少token,质量提升+9.2%;Fable达到+14.1%的最高质量。GPT-5.5系列模型(+6.2%至+8.6%)以更少token超越前沿。Gemini-3.5 Flash消耗token最多但效果不佳,Grok Build 0.1消耗20K+ token却出现负提升。

原文 · lmarena.ai

[Token efficiency in Agent Arena] Agent Arena measures agent performance across a range of real-worl...

[Token efficiency in Agent Arena] Agent Arena measures agent performance across a range of real-world tasks from our global community. Models get search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide deck, researching the web, building apps, and analyzing documents. More tokens in inference time may lead to better quality, but not always. We compare different models performance vs token consumption. The distance from the trend line is the interesting part. This chart maps every model in the Agent Arena: - Two axes: how much it improves real agentic work (improvement over avg, vertical) vs. how many tokens it burns to get there (median output, horizontal) - Diagonal line: shows the general trend. Output-token is generally linked to performance improvement. Key insights: - Opus consumes less tokens per session for better quality, Fable is the highest quality of all, reaching +14.1% compared with +9.2% for Opus 4.8 Thinking at similar token use. - All three GPT-5.5 models sit above the token efficiency frontier, ranging from +6.2% to +8.6% while using fewer tokens than the leading Claude models. - GLM-5.2 reaches +5.1% and sits close to the predictable trend line. - Tokens ≠ payoff. Gemini-3.5 Flash spends the most tokens, but much lower than frontier, and Grok Build 0.1 burns 20K+ for negative net improvement. Both far below the line. - Bottom-left is low-spend, low-gain: Grok-4.3, Nemotron 3 Ultra, and Gemma-4 31B sit where fewer tokens line up with weaker results. 💬 3 🔄 1 ❤️ 13 👀 2925 📊 5 ⚡