VulcanBench实测:Qwen3.8-Max慢且贵,要准确率选Grok 4.5、要性价比选DeepSeek V4-Flash

如果你想优化准确性,Grok 4.5 是首选。如果你想优化每美元准确性,DeepSeek V4-Flash 很难打败。👍 完成简单任务用 Grok 4.5/DeepSeek V4-Flash 挑战...

精选理由

VulcanBench 实测:Qwen3.8-Max 跑一趟 126 美元,DeepSeek V4-Flash 只要 13.6 美元,便宜 10 倍还更快。选性价比就 DeepSeek,选准确率就 Grok 4.5。

AI 摘要

VulcanBench 用 23 个真实软件工程任务评测模型,Qwen3.8-Max 跑完全套成本 126.25 美元,DeepSeek V4-Flash 仅需 13.60 美元。Qwen3.8-Max 每个任务耗时 20-25 分钟,是测试中最慢的模型,固定预算下最佳设置只排中游,默认设置垫底。六个原本能解决的任务贡献了 26 分降幅的 83%,其中三个得分直接归零。在准确率上 Grok 4.5 领先,按每美元准确率 DeepSeek V4-Flash 很难被击败。

原文 · Geek

如果你想优化准确性,Grok 4.5 是首选。如果你想优化每美元准确性,DeepSeek V4-Flash 很难打败。👍 完成简单任务用 Grok 4.5/DeepSeek V4-Flash 挑战...

如果你想优化准确性,Grok 4.5 是首选。如果你想优化每美元准确性,DeepSeek V4-Flash 很难打败。👍 完成简单任务用 Grok 4.5/DeepSeek V4-Flash 挑战困难任务用 Claude Fable 5/GPT-5.6 Sol 大模型不存在中间档的,谢谢,还有问题吗。 Morgan @morganlinton Okay, the @VulcanBench results for Qwen3.8-Max are in, and it is not what I expected. First, for anyone new to VulcanBench, here's a quick TL;DR on the eval suite: 23 frontier-hard software engineering tasks taken from real merged OSS PRs, run in a Docker sandbox, 3 runs per task across all three of its effort levels. No puzzles, no random abstract stuff, all real things engineering teams would do with these models. It looks like Qwen3.8-Max has a major overthinking problem, it uses a LOT of tokens and is very slow, period, no other way to see it. My cost to run this benchmark was $126.25, to run the exact same eval suite with DeepSeek V4-Flash was only $13.60. This makes Qwen3.8-Max an insanely expensive model. The tasks Qwen genuinely can't solve fail at every effort level, extra reasoning didn't help. The regression is almost all in work it already handles: six tasks that low solves every single time account for 83% of the 26-point drop, three of them collapsing to zero. It's not losing the hard problems. It's losing the ones it already knows how to do. Since Qwen3.8-Max hit a lot of wall clock budget caps, I thought I'd share more about this. - VulcanBench caps both steps (50–200) and wall clock (5–60 min), each scaled by repo size. - This is aligned with how comparable harnesses bound agents, DeepSWE caps rollouts at 100 environment steps, sitting right inside my step range; Terminal-Bench enforces a per-task wall clock; SWE-bench Verified scaffolds typically allow 20–60 min per instance with 250–350 step limits. - Every model on my chart gets the identical budget, and Qwen is the slowest model I've tested at 20–25 min/task. Soooo... Alibaba positions Qwen3.8-Max as trailing only Claude Fable 5. But on the kind of real coding work engineering teams would actually throw at it, under a fixed budget, its best setting lands mid-pack and its default lands last, so common. If you want to optimize for accuracy, Grok 4.5 is the move. If you want accuracy per dollar, DeepSeek V4-Flash is hard to beat, heck it's 10× cheaper than Qwen and you get higher accuracy. Qwen just isn't in the game at this point, this is not a model I could see engineering teams using for daily coding work. 🔗 View Quoted Tweet 💬 3 🔄 0 ❤️ 2 👀 829 📊 2 ⚡