看Opus 5和GPT-5.6-Sol在同一基准上的表现,一个越强越费token,一个越强越省,挺有意思的对比。
Agent Arena追踪模型完成真实任务所需的token数量。Opus系列从4.7到5,性能持续提升,但token消耗从约8.5k增至约21k,含推理和输出token。相反,GPT-5.5到GPT-5.6-Sol的token用量从约10k降至约8k,同时净改进提升约1.5pp。两组模型在效率与性能上呈现相反趋势。
Agent Arena tracks the number of tokens a model takes to complete real-world tasks. We see the perf...
Agent Arena tracks the number of tokens a model takes to complete real-world tasks. We see the performance of Opus-series models has improved significantly (Opus 4.7 to 4.8 to 5), and at the same time token usage has also increased substantially (~8.5k for Opus 4.7 up to ~21k for Opus 5), when counting both reasoning and output tokens. The opposite trend appears for top GPT models. Between GPT-5.5 and GPT-5.6-Sol, we see the token usage declined from ~10k to ~8k, despite a notable performance gain of ~1.5pp of Net Improvement. 💬 8 🔄 7 ❤️ 44 👀 4865 📊 13 ⚡