Claude Opus 5.5 登顶 FrontierCode 1.1 Extended 榜首
On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and qua...
Cognition 自家的 FrontierCode 榜更新了,Claude Opus 5.5 拿 65.3% 超过 Fable 5,成本还更低,写代码的可以看看。
Cognition 发布的 FrontierCode 1.1 是评测真实工程任务的基准,会考察代码合并可行性与质量。在该基准的 Extended 集上,Claude Opus 5.5 拿到 65.3%,取代 Fable 5 成为第一。Cognition 提到其成本远低于原榜首模型。
On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and qua...
On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and quality, Opus 5.5 scores 65.3% on Extended — taking the leading spot from Fable 5 at a fraction of the cost. Read more: devin.ai/blog/claude-op… 💬 1 🔄 0 ❤️ 2 👀 283 📊 1 ⚡
- 宝玉16:22原文