ProximalHQ 推出了超长编码基准 v2,Claude Fable 5.1 在测试中大幅领先其他模型。
ProximalHQ 发布 FrontierSWE v2 超长编码基准测试,更新了任务套件和方法论。测试显示前沿模型间存在巨大性能差距,Claude Fable 5.1 领先优势明显,超过 25 个百分点。该基准专注于超长视野编码任务,是评估前沿模型能力的重要工具。
Brilliant effort worth checking out. Ultra-long horizon coding tasks are where frontier models lik...
Brilliant effort worth checking out. Ultra-long horizon coding tasks are where frontier models like Fable 5.1 will shine. But that's a crazy gap (over ~25 percentage points). What I think could be interesting is seeing results for a mixture of agents like what Cursor did. Proximal @ProximalHQ We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large performance gaps between frontier models, with Claude Fable 5.1 leading by a wide margin 🔗 View Quoted Tweet 💬 4 🔄 0 ❤️ 7 👀 1348 📊 3 ⚡