METR 编码基准饱和？Cognition 发布 FrontierCode 新评测，Claude Opus 4.8 仅 13.4%

精选理由

做 AI 编程评估或关注模型实际能力的开发者，这个新基准直接戳中了当前模型的软肋——代码能跑但不可维护，值得看看你的模型能拿几分。

AI 摘要

Gary Marcus 发推指出 METR 的编码基准已饱和，但 Cognition 随即推出更难的 FrontierCode 评测，最高分仅 13.4%。该评测由顶级开源维护者花费 40+ 小时设计，首次衡量代码是否可合并维护，而非仅功能正确。这揭示了当前模型在编写可维护代码方面的严重不足，为 AI 编程能力评估设立了新标准。

AI 翻译 · 中文

Gary MarcusOh my God! @METR_Evals ’s coding benchmarks are saturated! 🤯 Mythos broke the METR graph 🤯 4 weeks later, out comes a new coding task, this time from @cognition : “FrontierCode Diamond remains unsaturated: the best per…

shao__meng06-09 01:01原文
rohanpaul_ai06-09 12:32原文
lmarena.ai06-09 23:56原文

查看原推