Google 提出正则化方法 RRSI,抑制自改进智能体的基准记忆化
Google researchers find a way to keep self-improving AI agents from memorizing their tests
Google 搞了个叫 RRSI 的正则化方法,专治自改进 Agent 背题,新基准上多拿 4.7 分还省三成 token。
Google 研究人员发现自改进 AI 智能体常把训练中的基准任务记下来,导致在新任务上的收益缩水。他们提出 RRSI 方法,通过正则化抑制这种记忆化效应。在未见过的基准上,RRSI 将分数提升最多 4.7 分,同时比未正则化版本节省约 30% 的 token。
Google researchers find a way to keep self-improving AI agents from memorizing their tests
Self-improving AI agents tend to memorize their test tasks, so their gains shrink or disappear on new ones. RRSI, a new method from Google researchers, reins in this effect and lifts scores on unseen benchmarks by up to 4.7 points while using about 30 percent fewer tokens than an unregularized version. The article Google researchers find a way to keep self-improving AI agents from memorizing their tests appeared first on The Decoder .