这项研究揭示了模型规模如何影响神经元的可解释性和专门化,对理解大模型内部机制和设计更高效架构的AI研究者有直接参考价值,建议关注其缩放定律的实践意义。
该研究探讨了神经网络中神经元群体是否随模型规模可预测地演化,扩展了缩放定律至损失等宏观可观测指标之外。通过分析高达30B参数的语言模型和5B参数的视觉模型,发现Rosetta神经元(跨独立训练模型激活模式相似的神经元)数量随规模呈亚线性幂律增长,但占总神经元比例缩小。研究还观察到“神经元极化效应”:Rosetta神经元随规模增加变得更选择性、更单语义,而非Rosetta神经元则保持较低选择性。一个平衡特征效用与有限神经元容量的分析模型解释了这种亚线性缩放和极化效应。结果表明存在可解释的、共享的神经元级结构缩放定律,将模型大小与神经元普遍性、选择性和专门化的系统性变化联系起来。
Neuron Populations Exhibit Divergent Selectivity with Scale
We investigate whether neuron populations within neural networks evolve predictably with scale, extending scaling laws beyond macroscopic observables such as loss. To probe this question, we study Rosetta Neurons, a previously characterized class of neurons whose activation patterns are similar across independently trained models (Dravid et al., 2023). In separate analyses of language models up to 30B parameters and vision models up to 5B parameters, we observe that the population of Rosetta Neurons follows a sublinear power law in model size, growing in absolute number but occupying a shrinking fraction of the total neuron count. We further observe a Neuron Polarization Effect: Rosetta Neurons become more selective and increasingly monosemantic with scale, separating from a growing non-Rosetta population that remains less selective. An analytical model balancing feature utility against limited neuron capacity explains the sublinear power-law scaling and this polarization effect. Finally, we find that Rosetta Neurons become more domain-specialized with scale and illustrate their selectivity through a targeted data-filtering case study for continued pretraining. Our results point to a scaling law for interpretable, shared neuron-level structure, linking model size to systematic changes in neuron universality, selectivity, and specialization.