原子特征理论:SAE 扩大规模时会稳定复现同一批特征
A Testable Theory of Atomic Features
一篇 arXiv 论文,用实验证明 SAE 从 512 扩到 131,072 后特征稳定不分裂,做可解释性研究的可以看看。
论文提出语言模型表示中存在原子特征的理论,并给出可测试的“恢复原理”:规模递增的稀疏字典(如 SAE)会依次恢复训练数据中最普遍的原子。由此推出三个可验证预测:小 SAE 的特征会被所有更大的 SAE 保留、不同数据训练的 SAE 共享双方都普遍的特征、足够大的 SAE 同时恢复父特征与子特征。研究者在 512 到 131,072 规模的 SAE 上、基于两个大型 embedding 模型验证了这些预测,推翻了 SAE 特征随规模增大而“分裂”不稳定的传统看法。
A Testable Theory of Atomic Features
We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training data. This "recovery principle" yields three testable predictions: many features in small SAEs are shared by all larger SAEs, SAEs trained on different data share features prevalent in both, and sufficiently large SAEs recover both parent and child features. In contrast to conventional wisdom that SAE features are unstable and "split" as size increases, we find that these predictions hold on SAEs of sizes ranging from 512 to 131,072 trained on two large embedding models. From a theoretical perspective, our results suggest the promise of a scientific theory of representations based on atomic features. Practically, our results suggest the promise of scaling SAEs.