SteerCheck: Activation-Steering审计中的归因特异性和对齐泄漏

SteerCheck: Attribution Specificity and Alignment Leakage in Activation-Steering Audits

精选理由

SteerCheck是一个强大的审计工具,可以帮助研究人员识别和评估激活引导中的对齐泄漏问题,对于确保AI模型的可靠性和透明度非常有用。

AI 摘要

SteerCheck是一种预注册的归因审计工具,用于匹配离目标KL和分离均值、保护尾部、极性、转移和语义主张。研究发现,对齐泄漏与签名的余弦值强相关,25.3%的样本余弦值超过0.5,所有超过观察到的平均效果的样本余弦值均超过0.8。SteerCheck使得这些条件性和混合结论可审计。

原文 · arXiv: DeepSeek

SteerCheck: Attribution Specificity and Alignment Leakage in Activation-Steering Audits

Activation steering can change behaviour without establishing that the effect is specific to the intended concept. We introduce SteerCheck, a preregistered attribution audit that matches off-target KL and separates mean, protected-tail, polarity, transfer, and semantic claims. Exact replay of 960 Qwen3-14B interventions reveals complementary limits of common controls: isotropic directions occupy a narrow near-orthogonal region, whereas sign-randomized same-construction directions often retain substantial target alignment. Effect is strongly associated with signed cosine within the sign-randomized family ($ρ=.94$); $25.3\%$ of its draws exceed cosine $.5$, and every draw exceeding the observed mean effect has cosine above $.80$. This alignment leakage does not by itself invalidate a conditional randomization test; it limits what the comparator can distinguish and motivates reporting exchangeability assumptions, a construction diagnostic $A$, and the empirical cosine distribution. The primary Qwen complete gate remains negative because the protected tail fails all families. On independent data, continuous margin transfers only in Qwen and accuracy transfers in no selected cell. Prospectively registered language controls pass the complete gate in Qwen and DeepSeek, while a passing DeepSeek detox comparator rules out categorical separation; all nominal passes are sensitive to $Γ=1.10$. Frozen three-rater open-generation evaluation supports factual correction in DeepSeek but not Qwen; the automatic judge fails calibration (macro-F1 $.562$), so null-wide semantic results remain descriptive. SteerCheck makes these conditional and mixed conclusions auditable.