机制研究:语言模型如何检测并定位内部激活扰动
A mechanistic study of language model introspection
给想让 LLM 报告内部状态靠谱一点的人:这篇用 attention head 层面拆解了模型怎么发现和定位被注入的激活扰动。
一项 arXiv 论文用固定输入文本的受控任务研究 LLM 内省:向 10 个 token 位置之一注入 concept vector 或不做干预,让模型报告扰动位置。跨三个模型家族,论文找到两组注意力头:中间层的 gate heads 决定是否报告变化,后续层的 router heads 帮助选择报告的位置。对 gate heads 的干预可以在 router heads 已提供位置信息时抑制位置报告。定位更准确的 concept vector 会在 gate heads 中产生更强的注意力分数与输出响应,对应 QK、OV 计算中 key 与 value 变化更一致。
A mechanistic study of language model introspection
Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention occurred. How do models detect and localize such internal changes? We study this question using a controlled task that keeps the input text fixed. We either inject a concept vector into the hidden state at one of ten token positions or apply no intervention. The model is asked to identify the perturbed position or report that no intervention occurred. Across three model families, we identify two small groups of attention heads with distinct roles in introspective reporting. Middle-layer gate heads influence whether the model reports a change, while router heads in a later layer help select the position to report. Interventions on gate heads can suppress position reports even when router heads supply location information. We further examine why reporting accuracy varies across concepts. Concept vectors that are localized more accurately produce stronger attention-score and output responses in gate heads, which is associated with better alignment of the induced key and value changes in their QK and OV computations. Together, these findings identify attention-head mechanisms supporting introspective detection and localization.