OpenAI 提出编码模型评估需更严更公平更可信

As coding models improve, evals need to become harder, fairer, and more trustworthy. Better benchma...

精选理由

OpenAI 自己说现在的编码测试太简单了,得加点难度和公平性。搞 AI 写代码的可以看看基准怎么改。

AI 摘要

OpenAI 在推特上指出,随着编码模型不断进步,现有的评估基准已不足以衡量真实进展。他们强调需要设计更难的测试、更公平的对比和更可靠的结果。这条观点引发社区对当前基准如 SWE-bench 等是否仍有效的讨论。

原文 · OpenAI

As coding models improve, evals need to become harder, fairer, and more trustworthy. Better benchma...

As coding models improve, evals need to become harder, fairer, and more trustworthy. Better benchmarks help the field understand real progress and where the frontier is moving. 💬 4 🔄 2 ❤️ 169 👀 20344 📊 16 ⚡