DeepSeek-V4.1-Flash (Max) 发布,性能成本比提升 4.87%,成本仅 0.07 美元/任务
DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net impro...
DeepSeek 发布了新模型,性能提升 4.87%,成本只有 0.07 美元,比很多模型便宜很多。
DeepSeek-V4.1-Flash (Max) 在 Agent Arena 基准测试中实现了 +4.87% 的净提升,成本为每任务 0.07 美元。在顶级开源模型中,它以最低的成本达到这一性能水平。与 Hy4 preview 相比,它保留了 98% 的性能提升,但成本降低了 73%。与 Kimi K3 (Max) 相比,它保留了 76% 的性能,但成本降低了 92%。
DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net impro...
DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net improvement at $0.07 cost per median task, it’s reshaped the Pareto frontier for Agent Arena! Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost. For comparison, it retains: - 98% of Hy4 preview’s net improvement, at 73% lower cost - 76% of Kimi K3 (Max)’s performance, at 92% lower cost. Against models as powerful as Fable 5 or stronger, DeepSeek-V4.1-Flash (Max) retains 35–54% of their net improvement at 97–99% lower cost. Those top models cost 37–76× more per task. Net improvement over Arena baseline | Median cost/task: - Claude Fable 5.1 (Max): +13.90% | $4.54 - GPT 6 Astra (Max): +11.90% | $4.09 - Claude Opus 5 (Max): +11.09% | $3.52 - Claude Opus 5 (High): +10.49% | $2.24 - Claude Fable 5 (High): +9.03% | $2.19 - Claude Opus 4.8 (High): +7.75% | $1.36 - GPT 5.6 Sol (xHigh): +7.40% | $1.09 - Kimi K3 (Max): +6.39% | $0.77 - Hy4 preview: +4.96% | $0.22 - DeepSeek-V4.1-Flash (Max): +4.87% | $0.06 With this release, GPT-5.6 Luna (xHigh), GLM-5.3-Flash, and DeepSeek-V4-Flash fell off the Pareto frontier for Agent Arena. Congrats again to the @deepseek_ai team on this release! Your browser does not support the video tag. 🔗 View on Twitter Arena.ai @arena Exciting news: DeepSeek-V4.1-Flash (Max) by @deepseek_ai just landed in Agent Arena at #3 among open models! With +4.87% net improvement and a median cost per task of $0.07 it reshaped the Pareto frontier. Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median cost per task. Its +4.87% net improvement is within 0.09 percentage points of Hy4 preview (ranked #2 ) at 68% lower cost, and within 1.52 percentage points of Kimi K3 (Max) (ranked #1 ) at 91% lower cost. - Kimi K3 (Max): +6.39% | $0.77/task - Hy4 preview: +4.96% | $0.22/task - DeepSeek-V4.1-Flash (Max): +4.87% | $0.07/task See the full Pareto Frontier below. DeepSeek-V4.1-Flash (Max) is ranked #12 overall, and by signal landed #4 Confirmed Success with +13.75%! Congrats to the @deepseek_ai team on this release! 🔗 View Quoted Tweet 💬 3 🔄 8 ❤️ 101 👀 7536 📊 13 ⚡