AgentBug-Smith:把真实 bug 报告变成可运行测试,揭示 coding agent 修复能力不足
一项来自美中多所实验室的研究,把真实 agent bug 变成可跑的测试集。最好的 agent 只修好 9%,很扎心但很实用。
AgentBug-Smith 能把 GitHub 上真实的 bug 报告自动转换成可运行的测试,构建出一个包含 200 个 bug 且持续增长的基准。这些 bug 来自 agent 自身代码(工具调用、记忆、提示词),依赖真实模型调用,难以复现和测试。3 个 coding agent 中表现最好的只修复了 9% 的 bug,远低于普通软件 bug 约 40% 的修复率。加入一份来自过往修复经验的经验指南后,一个 agent 在 79 个未见过的 bug 上从修复 1 个提升到修复 6 个。
Self-improving AI agents will need to fix their own code.
And this paper from top US+China labs, shows coding agents miss most such bugs but improve with lessons from past fixes.
that real bugs in agent harnesses, can be automatically turned into a growing set of runnable tests.
An agent's own code is everything around the model: tool calls, memory, and prompts. Its bugs depend on live model calls, which makes them hard to recreate and test.
So the researchers built AgentBug-Smith, which turns real GitHub bug reports into runnable tests. The result is a 200-bug benchmark that keeps growing.
The best of 3 coding agents fixed just 9% of those bugs, versus about 40% reported on regular software bugs. A short guide of lessons from past fixes lifted an agent from 1 to 6 correct fixes on 79 unseen bugs.
Before trusting a coding agent with your agent's code, try it on bugs you've already fixed.