这篇论文展示了如何利用稀疏自编码器在ν中微子基础模型中寻找可解释的潜在表示,这对于理解模型的内部工作原理和设计下游任务非常有帮助。与传统的方向头相比,这种新的不确定性头在预测角重建误差方面表现出色。
本研究首次将稀疏自编码器机制可解释性应用于粒子物理学。通过研究在IceCube数据上预训练并针对方向重建微调的ν中微子基础模型,我们使用严格的验证协议(包括保留测试、匹配干扰控制和独立字典训练的复制)识别了模型表示中的验证物理概念图。因果干预表明,方向头几乎没有利用这个图。受此未充分利用的信息启发,我们在同一事件级表示上训练了一个不确定性头,以预测模型的角重建误差。与方向头不同,它因果依赖于图中的质量和亮度特征。在20%的选择效率下,这个可解释估计器将中值角分辨率从20.2°提高到3.2°。这些结果表明,机制可解释性可以揭示模型内部表示中编码的学到的潜在物理,并有助于设计利用它的下游任务。
Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At $20\%$ selection efficiency, this interpretable estimator improves the median angular resolution from $20.2^\circ$ to $3.2^\circ$. These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.