精选理由
Opus 5 在代码工程基准上接近顶级模型,但价格只要一半,做调试分析特别强。
FrontierCode 1.1 是 Cognition 发布的评估真实工程任务的基准,考核代码合并性和质量。Opus 5 在该基准上取得 63.6% 的分数,扩展测试通过率 69.6%,性能接近 Fable 5 而成本仅为后者的一半。在 Devin 环境中,Opus 5 在困难调试和根因分析任务上表现突出。
原文 · Cognition
On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and qua...
On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and quality, Opus 5 scores 63.6% with a 69.6% pass rate on Extended — approaching Fable 5 at half the cost. Within Devin, it shows particular strength on difficult debugging and root-cause analysis tasks. devin.ai/blog/claude-op… 💬 1 🔄 1 ❤️ 1 👀 401 📊 1 ⚡