给AI科学应用开发者的新视角:AlphaFold2的权重里藏着比静态预测更丰富的构象信息,用高斯卷积就能读出蛋白质折叠动力学线索。
研究提出对AlphaFold2的9300万参数进行直接分析,通过高斯卷积平滑Evoformer权重张量,发现训练后的模型产生物理结构的构象图谱。泛素扰动下原生接触断裂顺序与数十年折叠实验一致。五个独立训练的KaiB模型均未恢复替代折叠,而α-突触核蛋白产生五种不同但连贯的图谱。随机噪声控制证实构象源于学习而非噪音,该涌现特性非显式训练目标。
Neural spectroscopy of AlphaFold2 reveals encoded protein conformational landscapes
AlphaFold2's 93 million parameters, shaped by the evolutionary record of protein structure encoded in the Protein Data Bank and in sequence alignments, are conventionally treated only as machinery for converting sequence to structure. We propose they are also a scientific object that can be analyzed directly: a learned encoding of protein conformational organization that can be probed and characterized. By smoothing the Evoformer's weight tensors with a Gaussian convolution and scaling the result, we show that the trained model produces physically structured conformational landscapes. Under perturbation, ubiquitin's native contacts break in the order established by decades of folding experiments. For KaiB, five independently trained models agree that the alternative fold is not recovered under perturbation. For alpha-synuclein, five models produce five different but coherent landscapes, mapping where the training signal has determined the representation and where it has not. Matched-power noise controls confirm that random corruption of equal magnitude produces debris, not conformations. The model learned to predict static structures; the conformational organization visible under perturbation was not an explicit training target, suggesting it emerged as a byproduct of that objective. AlphaFold2's weights appear to encode structural constraints, shaped by evolutionary and structural training data, that extend beyond what unperturbed inference reveals. We call the approach of reading them neural spectroscopy, and Scaled Gaussian Convolution one such protocol.