Sonnet 5 在 FrontierCode 基准得分 53.8% 超越 Opus 4.8

On FrontierCode (Extended), our benchmark for real-world engineering tasks that grades mergeability ...

精选理由

Cognition 用 FrontierCode 实测了 Sonnet 5 的工程能力,它比 Opus 4.8 更强,值得关注。

AI 摘要

Cognition 发布了 FrontierCode (Extended) 基准,用于评估真实世界编程任务的可合并性和质量。Claude Sonnet 5 在该基准上得分 53.8%,通过率为 57.6%。该成绩高于 Opus 4.8 的表现,但 Cognition 表示随着后续基准调整,相对排名可能发生变化。

原文 · Cognition

On FrontierCode (Extended), our benchmark for real-world engineering tasks that grades mergeability ...

On FrontierCode (Extended), our benchmark for real-world engineering tasks that grades mergeability and quality, Sonnet 5 scores 53.8% and has a 57.6% pass rate (higher than Opus 4.8). Note: These relative rankings may change slightly with coming adjustments to FrontierCode. devin.ai/blog/claude-so… 💬 1 🔄 3 ❤️ 27 👀 3497 📊 4 ⚡