论文精选72°

LLM智能体记忆不可靠:反复重写反而更糟,清华等团队新研究

New Illinois+ Tsinghua University and other labs s…

精选理由

做智能体系统或记忆管理的开发者,这篇论文戳中了记忆重写的致命缺陷——原始经验比精炼总结更可靠,看完你会重新思考记忆存储策略。

AI 摘要

伊利诺伊大学、清华大学等机构联合研究发现,LLM智能体在反复重写自身记忆时,记忆可靠性会下降。许多智能体系统通过让LLM将原始经验压缩成整洁的书面总结来存储记忆,但论文指出,这种反复重写会逐渐损害记忆。实验表明,原始经验(即实际尝试和解决方案)往往比精炼的总结更有用。例如,GPT-5.4在无记忆情况下能100%解决ARC-AGI谜题,但使用基于正确解构建的记忆后,流式更新使成功率降至约54%。失败原因包括错误分组、过度泛化和过拟合,导致记忆丢失细节、混淆任务类型或学习到仅适用于狭窄案例的规则。论文建议,智能体记忆不应自动将每次经验重写为摘要,保留原始证据并偶尔进行总结效果更好。

原文 · rohanpaul_ai

New Illinois+ Tsinghua University and other labs s…

New Illinois+ Tsinghua University and other labs study finds that LLM agents still have unreliable memory and that it can get worse when they keep rewriting their own memories.

LLM agents can learn from experience, but their rewritten memories often become unreliable.

The problem is that many agent systems store past work by asking an LLM to compress messy experience into neat written lessons.

That sounds useful because the agent should remember what worked before, but the paper finds that repeated rewriting slowly damages the memory.

The core idea is that raw episodes, meaning the actual past attempts and solutions, often stay more useful than the polished lessons made from them.

The authors tested this across tasks like web shopping, simulated worlds, app use, and ARC-style puzzle problems where they could control the correct solutions.

The sharpest result is that GPT-5.4 solved 100% of a small ARC-AGI set with no memory, but after memory was built from correct solutions, streaming updates dropped it to about 54%.

The failures came from bad grouping, overbroad lessons, and overfitting, so the memory forgot details, mixed up task types, or learned rules that only worked on narrow examples.

The big deal is that agent memory should not automatically rewrite every experience into a summary, because keeping raw evidence and only sometimes making summaries worked better.

The paper is really proposing that agent memory should treat raw past episodes as important evidence, not as disposable notes to summarize away.

----

Paper Link – arxiv. org/abs/2605.12978

Paper Title: "Useful Memories Become Faulty When Continuously Updated by LLMs"