论文多源确认

自我进化 LLM 系统何时该停:arXiv 论文提出即插即用停止准则

When Is Enough Enough in Self-Evolving LLM Systems?

精选理由

做自我进化 agent 的可以看看:不用改算法,套个统计检验就能让系统自己判断什么时候停止迭代,省 91.6% 的 token 还不掉点。

一篇 arXiv 论文研究自我进化 LLM 系统的停止时机与输出选择问题。作者把停止判定建模为在线序贯检验,用每次迭代已产生的逐项配对评估结果构建 anytime-valid 重启检测器,再用变点估计从历史候选中选一个更早的产物输出。该方案即插即用,无需修改底层自我进化算法。在 2 个自我进化框架、3 个 LLM 模型家族、5 个基准上的实验显示,SearchQA 上 SkillOpt 与 DeepSeek V4 Flash 组合在第 4 轮即停止(原预算 40 轮),token 用量减少 91.6%,未见测试集准确率 82.43% 对比跑满预算的 82.00%。

原文 · arXiv: DeepSeek

When Is Enough Enough in Self-Evolving LLM Systems?

Self-evolving large language model (LLM) systems repeatedly propose, evaluate, and incorporate updates to prompts, skills, or other persistent artifacts. Despite their growing effectiveness, these systems typically operate under a predetermined iteration or compute budget, without a principled criterion to determine when further evolution is no longer worthwhile. This can lead to two undesirable consequences: unnecessary computation after performance has saturated and the risk of returning late updates that overfit or exploit the evaluation signal. These issues motivate us to study two fundamental questions: when should a self-evolving system stop, and what should it output once it stops? We address the first by formulating an online sequential testing problem and constructing an anytime-valid restart detector using the per-item paired evaluation outcomes already produced by self-evolving LLM systems. We address the second by formulating a change-point estimation problem and using the estimated transition to select an earlier artifact for output. The resulting procedure is plug-and-play and requires no modification of the underlying self-evolving algorithms. Across two self-evolving frameworks, three LLM model families, and five benchmarks, our method substantially reduces computation costs while maintaining comparable unseen-test performance. For example, on SearchQA with SkillOpt and DeepSeek V4 Flash, our method stops at round 4 rather than the full budget of 40, reducing token usage by 91.6% while achieving 82.43% unseen-test accuracy versus 82.00% under the full-budget run.