Agent Arena 分析 20840 条轨迹:编码智能体差评首因是代码跑不通
What can user praise and complaints tell us about coding agents themselves? We analyzed 20,840 trac...
Arena 拆了 20840 条编码智能体反馈:69% 差评是代码跑不通,Fable 5.1 比 Astra 更少被吐槽。
Arena 团队分析了 Agent Arena: Code 平台上 29 个模型的 20840 条用户反馈轨迹,用直接的正负评价来衡量编码智能体的实际表现。投诉中 69.1% 指向代码无法运行,其后是输出不完整(27.1%)、收尾与可用性差(27.1%)和忽视指令(20.3%)。Fable 5.1 的“slop”与设计类投诉占 3.9%,低于 Astra 的 5.7%,“flaky”类投诉占 3.1%,也低于 Astra 的 5.0%。29 个模型的数据还显示,各实验室的新一代前沿模型获得的正反馈占比普遍更高。
What can user praise and complaints tell us about coding agents themselves? We analyzed 20,840 trac...
What can user praise and complaints tell us about coding agents themselves? We analyzed 20,840 traces from Agent Arena: Code across 29 models, and found that direct feedback unlocks novel opportunities in tracking the frontier. Some topline findings: - Today’s frontier models receive markedly more positive feedback. Overall, newer models across labs tend to show a more positive feedback balance. - Broken code is still the biggest driver of complaints, sloppy behavior comes second. 69.1% of complaints point to code not working. The next themes are incomplete output (27.1%), weak finish and usability (27.1%), and ignored instructions (20.3%). - Models share common weaknesses, but differ in how they disappoint. Fable 5.1 attracts fewer “slop” and design complaints than Astra (3.9% versus 5.7% of sampled traces), and fewer complaints about being “flaky” (3.1% versus 5.0%). Read more by diving into the full article from @DawidGalarowicz below. Arena.ai @arena x.com/i/article/2101… 🔗 View Quoted Tweet 💬 4 🔄 1 ❤️ 45 👀 4431 📊 7 ⚡