论文

AnchorLoop:引入历史参考的自进化智能体训练方法

Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents

精选理由

MIT团队新论文AnchorLoop,让智能体通过历史参考自我进化,数学推理提升2.5%,比传统方法持续改进。

研究人员提出AnchorLoop方法,通过冻结前一代执行器作为历史参考,改进自进化工具集成智能体的训练过程。该方法在13个推理基准测试中,将数学推理性能提升2.5%,一般推理任务提升2.8%。AnchorLoop为执行器提供跨参考优势,为课程设计提供基于一致性的参考,无需外部任务或答案监督。

原文 · arXiv cs.AI

Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents

Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop. A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning. However, relying solely on the current Executor for feedback has two limitations: group-relative advantages vanish under full consensus, while uncertainty-based curriculum rewards favor disagreement without showing whether the generated tasks support further learning. These limitations motivate an additional reference beyond the current Executor. We propose \textit{AnchorLoop}, which introduces a frozen copy of the previous iteration's Executor as a historical reference and reuses it on both sides of the training loop. For the Executor, the anchor provides a cross-reference advantage that evaluates current outputs against both current and historical majority answers. For the Curriculum, it provides an agreement-based reference based on differences in sampled majority agreement. Since the Executor and anchor have identical parameters during Curriculum training, this comparison serves as a proxy for task selection rather than evidence of inter-version improvement or correctness. Across 13 reasoning benchmarks, AnchorLoop improves over Agent0 by 2.5\% on mathematical reasoning and 2.8\% on general reasoning tasks. It also maintains higher effective-advantage variance and continues improving in later iterations as the unanchored baseline shows diminishing gains. These results demonstrate the benefit of introducing a lightweight historical reference into self-evolving tool-integrated agents without external task or answer supervision.