行业精选78°

OpenAI审计发现SWE-Bench Pro已饱和,建议停用

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer r...

精选理由

OpenAI发现SWE-Bench Pro这个编码基准有70%噪音上限,已经饱和了,不再适合用来衡量最强AI编码能力。

AI 摘要

OpenAI对SWE-Bench Pro进行了审计,发现该基准存在约70%的噪音上限,已无法可靠衡量前沿编码能力。该基准在业界广泛使用,但目前已被认为饱和。OpenAI决定撤回之前推荐研究社区使用该基准作为主要编码评估的建议。

图片来源 · OpenAI
原文 · OpenAI

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer r...

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find the eval to be saturated at a ~70% noise ceiling, and are retracting our previous recommendation that the research community use it as a leading coding eval. openai.com/index/separati… 💬 105 🔄 110 ❤️ 1670 👀 114561 📊 272 ⚡

OpenAI审计发现SWE-Bench Pro已饱和,建议停用 · AI 热点