论文精选

SelfSearch:无奖励搜索提升智能体自我改进

SelfSearch: Reward-Free Search for Self-Improving Agents

精选理由

SelfSearch让智能体通过自我修改记录提升能力,无需下游奖励信号,成本更低效果更好。

SelfSearch是一种无奖励搜索程序,智能体通过记录之前的自我改进经验来修改自身。在六种模型-基准设置中,SelfSearch将种群平均成功率提升超过初始智能体。在SWE-bench多语言测试中,智能体成功率提高5.0个百分点,同时降低38.5%的执行成本。使用仅4.03美元的搜索成本,DeepSeek V4 Flash在Terminal-Bench 2.1上达到82.0%的任务解决率。

原文 · arXiv: DeepSeek

SelfSearch: Reward-Free Search for Self-Improving Agents

Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}\% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{\$4.03} in search cost, it produces a harness that solves \textbf{82.0}\% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.