语言模型自修复机制研究:增益与反作用力
Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
这篇论文揭示了语言模型"自修复"的数学本质,解释了为什么移除组件后其他部分会补偿。
研究人员发现语言模型组件移除后的"自修复"现象实为预存增益的作用。研究团队在Gemma、Qwen、LLaMA和Mistral四个模型家族中验证了这一仿射定律,81个下游方向中有68个符合该规律。在GPT-2 Small的IOI电路中,10个可干预头部中有7个遵循此规律且均为反作用力。该研究揭示了传统消融方法在坐标轴上的未校准特性。
Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair
Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $λ$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(λ)=\mathrm{own}_r+γ_rλ$. The slope $γ_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $γ_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.