Cognition 用 FrontierCode 实测了 Sonnet 5 的工程能力,它比 Opus 4.8 更强,值得关注。
Cognition 发布了 FrontierCode (Extended) 基准,用于评估真实世界编程任务的可合并性和质量。Claude Sonnet 5 在该基准上得分 53.8%,通过率为 57.6%。该成绩高于 Opus 4.8 的表现,但 Cognition 表示随着后续基准调整,相对排名可能发生变化。
On FrontierCode (Extended), our benchmark for real-world engineering tasks that grades mergeability ...
On FrontierCode (Extended), our benchmark for real-world engineering tasks that grades mergeability and quality, Sonnet 5 scores 53.8% and has a 57.6% pass rate (higher than Opus 4.8). Note: These relative rankings may change slightly with coming adjustments to FrontierCode. devin.ai/blog/claude-so… 💬 1 🔄 3 ❤️ 27 👀 3497 📊 4 ⚡