这篇论文把AI自我改进的各种文献梳理得特别清楚,区分了“有界自优化”和真正的“递归自我改进”,还画出了验证强度层级,读完对AI安全风险理解更深。
该论文调查了2024-2026年间1250篇arXiv论文,沿两个轴分类:系统改进的对象(行为、策略、评估器、研究过程)和循环闭合程度。它将有界自优化(收敛、可评估、已成工业实践)与开放端递归自我改进(RSI)区分开,后者在所有可测维度上受制于基础、崩溃动态和计算约束。论文提出评估器设计空间,将信号从形式验证器(最强)到内在自我评估(最弱)排序,并证明自我改进强度与验证层级相关。失败模式如自确认循环、模型崩溃和多样性崩溃源于违反该层级,而保持人类在环中的“研究方向设定”瓶颈位于该层级顶端。
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data they generate, and, increasingly, conducting AI research itself. This literature is described under a vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") that conflates fundamentally different ambitions. We survey 1,250 arXiv papers (2024-2026) along two axes: what the system improves -- its behavior in deployment, its policy through training, its evaluator, or the research process itself -- and the degree of loop closure (human-in-the-loop to fully closed). The taxonomy separates bounded self-refinement -- convergent, evaluable, and already industrial practice -- from open-ended recursive self-improvement (RSI), which remains bounded by grounding requirements, collapse dynamics, and compute constraints on every measured axis. Its distinctive feature is a dedicated category for self-evaluation: every improvement loop is a claim that some signal can substitute for human judgment. We survey the evaluator design space -- judges, process reward models, verifiers, rubrics, meta-evaluation -- order the signals into a verification hierarchy from formal verifiers (strongest) to intrinsic self-assessment (weakest), and observe that demonstrated self-improvement strength tracks this hierarchy, that its failure modes (self-confirming loops, model collapse, diversity collapse) follow from its violations, and that the "research direction-setting" bottleneck keeping humans in the loop sits at the top of that hierarchy. We connect the technical literature to the theory of RSI limits and to the safety and governance questions raised by frontier-lab accounts of closing the loop, and identify governance-grade measurement of self-improvement as the field's most underpopulated niche.