GPT 5.5 称霸 SWE 基准，但 Opus 4.8 仍是 Vibe Coding 之王

精选理由

ViBench 填补了现有基准只测代码修复、不测完整应用创建的空白，做全栈原型或快速验证想法的开发者值得关注——Opus 4.8 可能才是你的性价比之选。

AI 摘要

尽管 GPT 5.5 在 SWE 基准测试中表现最佳，但 Opus 4.8 在端到端应用创建任务上仍保持价格与性能的双重优势。为此，团队推出了 ViBench——首个基于真实世界任务的应用创建基准测试。该基准旨在更准确地评估模型在实际开发场景中的表现，而非仅关注代码修复或补全。结果显示，Opus 4.8 在 Vibe Coding 场景下依然是最优选择。

AI 翻译 · 中文

Amjad MasadBenchmarks place GPT 5.5 as the best model on SWE, but is it the best at making apps end-to-end? Turns out Opus 4.8 continues to be the king of vibe coding on both price & performance. Introducing ViBench: the first …

宝玉06-04 17:26原文
Yangyi06-03 00:39原文

查看原推