论文

研究提出"虚假遗忘"机制:微调中"丢失"的知识可被恢复

When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting

精选理由

一篇挺有意思的机制分析:模型微调后"忘了"的东西其实还在,删掉一个方向就能捞回来,做微调的值得看看原理。

论文研究语言模型在微调时出现的 spurious forgetting(虚假遗忘):被"遗忘"的旧知识往往仍存储在模型中,甚至能在继续训练新事实后自行恢复。作者用最小联想记忆模型复现了这类动态,关键要素包括共享结构的 keys、集中的新值和归一化层。实验显示,减去表征中的共同偏移可消除 Transformer 上的遗忘崩溃,在预训练语言模型中移除每次权重更新中的单个方向即可恢复旧事实。结论是遗忘由可逆的访问丢失和不可逆的逐事实侵蚀两部分组成,只有后者才是真正的 catastrophic forgetting。

原文 · arXiv cs.LG

When Forgetting is not Catastrophic: On the Mechanics of Spurious Forgetting

Knowledge that a language model appears to forget during finetuning often remains stored and can be recovered, a phenomenon called spurious forgetting. Finetuning on new facts can even produce forgetting that undoes itself: recall of the old facts collapses, recovers as training continues on new facts alone, and only then erodes for good. We seek to understand when such forgetting is not catastrophic. A minimal associative memory reproduces these dynamics with three ingredients: keys with shared structure, concentrated new values, and normalization in the network. Finetuning moves all old representations along a common direction, hiding the old facts while preserving their relative geometry; normalization withdraws this shift once the new facts are learned, whereas fact-specific changes accumulate and cause the erosion. Moreover, subtracting the common shift eliminates the collapse in a Transformer trained on synthetic data, and removing a single direction from each weight update restores old facts in a pretrained language model. Forgetting thus combines a shared, reversible loss of access with a slow erosion of individual facts, and only the second is catastrophic. Which one dominates depends on whether the new data move old memories together or apart.