模型

DeepSWE测试显示DeepSeek V4.1-Flash输入token占比远高于输出

Many assume that agent spend goes toward output tokens. When we ran DeepSWE on Astra vs. DeepSeek V...

精选理由

DeepSWE测试发现DeepSeek V4.1-Flash的输入token消耗远高于输出,缓存命中率高,能大幅降低成本。

DeepSWE测试对比Astra和DeepSeek V4.1-Flash,输入token数量是输出的174倍。99.6%的请求是缓存命中,占60%的费用。最终任务质量相同,但成本从$6.52降至$0.43。

图片来源 · Fireworks AI
原文 · Fireworks AI

Many assume that agent spend goes toward output tokens. When we ran DeepSWE on Astra vs. DeepSeek V...

Many assume that agent spend goes toward output tokens. When we ran DeepSWE on Astra vs. DeepSeek V4.1-Flash, input tokens outnumbered output 174 to 1. 99.6% were cache hits. Those hits are 60% of the bill. Net result? Same quality. $0.43/task vs $6.52. x.com/i/article/2099… 💬 1 🔄 0 ❤️ 13 👀 756 📊 3 ⚡