论文

论文研究电路评估方法:可复现成功却解释不了模型错误

Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

精选理由

做可解释性的朋友看过来:这篇用 GPT-2 的 IOI 实验说明现有电路验证方法基本只看正确样本,错误复现率低到 11.4%,方法上值得反思。

一项机制可解释性研究提出,电路级解释不仅要复现模型的正确行为,也要能解释其错误。团队在 IOI、Docstring 及 Mechanistic Interpretability Benchmark 的六个模型-任务设置上做了测试。结果显示 GPT-2 small 的 IOI 手工电路与多个自动电路在正确样本上答案一致率达 97.3-99.5%,但错误样本上一致率只有 11.4-41.7%。案例研究发现补回被省略的注意力头后,错误复现率从 14.2% 升到 75.1%,正确一致率仅下降 0.41 个百分点。论文主张把错误复现作为电路解释的必要测试条件。

原文 · arXiv cs.LG

Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to recover the underlying mechanism of the model's behaviour by closely reproducing its successful decisions while failing to account for most of its errors. Such explanations should account for the model's particular errors as well as its successes. We evaluate this requirement by measuring exact answer agreement separately on model successes and failures, across circuit sizes and ablation settings, on IOI, Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark. We discover that many tested circuits closely replicate correct behaviour while missing most of the model's errors. On indirect object identification (IOI) for GPT-2 small, under mean ablation, the manual circuit and tested automated circuits, including one trained against the model's full output distribution, agree with the model on 97.3-99.5% of prompts it answers correctly but only 11.4-41.7% of errors. An IOI case study shows that lost errors are recoverable by restoring omitted attention-heads which raise error reproduction from 14.2% to 75.1% on a separate held-out set with 0.41 percentage point decrease on correct agreement, exceeding matched random extensions and scalar-biased control. Intervention traces show how omitted computations produce specific wrong answers for a reproducible subset of errors. In all, these findings show circuits can preserve task success without adequately explaining model's failures, and support exact error reproduction as a necessary, but not sufficient, test of circuit-based explanations of model behaviour.