RETRACE 从 CI 修复历史中提取可复用经验,提升智能体修复成功率
RETRACE: From Entangled Repair Histories to Reusable Experience for CI Repair
arXiv 上的新论文 RETRACE,教 AI 智能体从真实 CI 修复历史里学经验,mini-SWE-agent 成功率从 19.6% 涨到 31.9%,做智能体修复的可以看看。
论文提出 RETRACE 框架,用于从包含多次失败尝试和无关改动的 PR 历史中重建问题级修复经验。该框架结合从通过版本反向推理的 endpoint 视图和沿提交历史正向追踪的 development 视图,并用 CI 执行证据对齐两种视图。在 CI-REPAIR-BENCH 的 565 个 PR 级修复(覆盖 101 个仓库、12 类故障)上,RETRACE 将 mini-SWE-agent 配合 MiniMax-M2.5 的 Pass@1 从 19.6% 提到 31.9%,配合 DeepSeek-V4-Flash 从 23.3% 提到 32.8%,Codex 在匹配子集上从 15.5% 提到 27.5%。
RETRACE: From Entangled Repair Histories to Reusable Experience for CI Repair
Large language model (LLM) agents increasingly reuse prior experience, but most approaches assume that problems and solutions are already aligned. Software histories rarely provide this alignment: a pull request (PR) may contain multiple continuous integration (CI) problems, failed attempts, reverted edits, and unrelated changes, obscuring which changes resolve each problem. We present RETRACE, a framework for reconstructing problem-level repair experience from such histories. RETRACE combines an endpoint view that reasons backward from changes retained in the passing revision with a development view that traces repair evolution forward through commit history. CI execution evidence reconciles the two views, and the recovered experience is represented at three abstraction levels, from concrete fixes to transferable repair patterns. For new failures, RETRACE retrieves relevant problem-level experience to guide repair. On CI-REPAIR-BENCH, comprising 565 PR-level repairs from 101 repositories across 12 failure categories, RETRACE improves mini-SWE-agent Pass@1 from 19.6% to 31.9% with MiniMax-M2.5 and from 23.3% to 32.8% with DeepSeek-V4-Flash. On a matched subset, Codex improves from 15.5% to 27.5%. Combining both views consistently outperforms either alone, showing that recovering problem-change alignment enables historical CI repairs to serve as reusable repair experience.