稀疏自编码器特征集不稳定,不反映人类概念边界

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

精选理由

这篇论文用SAE潜变量集做相似度度量,发现它并不比稠密嵌入更贴近人类概念判断,搞可解释性的人值得看看这个反直觉结论。

AI 摘要

Shani等人2026年的研究显示LLM表征能大致还原人类类别边界,但无法捕捉细粒度典型性结构。本文用稀疏自编码器(SAE)活跃潜变量集的重叠度作为更可解释的相似度度量,重新审视该分析。在受控玩具模型中SAE潜变量集能恢复并集式组合结构,在自然文本中也能诱导语义连贯的邻域。但扩展到人类概念分析时,SAE激活集并未比稠密嵌入或残差流状态更忠实还原类别边界或典型性,而是追踪模型内部相似性结构。在受控语义修改下,人类概念变化判断与SAE活跃集变化存在显著不匹配,说明在理想化设置之外SAE特征不按简单词袋语义组合。

原文 · arXiv cs.LG

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts analysis to SAE set similarities, we find that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure. To probe this gap further, we study active latent sets under well-controlled semantic modifications, revealing a substantial mismatch between human judgements of conceptual change and change in the SAE active set. We interpret this as evidence that, outside idealised settings, SAE features do not compose via simple bag-of-features semantics.