自进化技能在未见任务上的泛化能力研究
Do Self-Evolving Skills Generalize to Held-Out Tasks?
这篇论文研究了AI自进化技能在未见任务上的泛化能力,提出了GSO方法,对AI开发者很有参考价值。
研究人员测试了五种自进化技能方法在六个基准上的表现。21个在训练任务上改进的技能中,5个在测试任务上保持了全部改进,13个保留了部分改进,3个没有保留任何改进。基于这些发现,研究人员提出了通用技能优化(GSO)方法,在所有六个基准上得分最高。
Do Self-Evolving Skills Generalize to Held-Out Tasks?
AI agents can externalize what they learn from past tasks into reusable \emph{skills}, such as procedures, checklists, code, or other executable artifacts, that can be retrieved and reused when solving new tasks. Self-evolving skill methods keep rewriting these skills after each round of practice on training tasks, and the skill is then used on new tasks of the same kind. We ask a question: does the improvement a skill shows on its training tasks carry over to new test tasks? We test five self-evolving methods and a one-shot skill on six benchmarks, with the same model, the same agent, and the same train/test split for every method. Of the 21 skills that improve on their training tasks, 5 keep all of that improvement on the test tasks, 13 keep part of it, and 3 keep none of it. No existing method is best everywhere. When we read the skills, the ones that carry over badly often fix details that should depend on the task, such as column names and output files, or turn a fix for one failure into a rule for every task. An LLM judge that reads the skill content can often see this: it ranks finished skills the same way the test results do in 86\% of pairs. But it predicts the effect of a single edit poorly, so edits still have to be tested by running them. Based on these findings, we describe Generalizable Skill Optimization (GSO), which keeps only a guide for writing skills and writes a new skill for each task; it scores highest on all six benchmarks.