做分子动力学模拟或材料计算的团队,终于有了一个能同时处理能量和力的主动学习框架,效率比委员会方法高得多,建议做MLIP微调的直接试试。
该研究提出一种基于分块特征空间后验方差筛选的线性扩展采集框架,避免候选集和训练集核矩阵的显式构建,可在数小时内筛选约20万结构。研究将神经正切核扩展到力感知场景,通过混合参数-坐标导数得到力NTK和联合能量-力NTK,为向量场预测提供自然相似性度量。在OC20数据集上,联合能量-力NTK在所有指标和分布划分下取得最低能量和力MAE及RMSE。在T1x、PMechDB和RGD基准测试中,力NTK方法在保持与基线竞争力同时,比基于委员会的方法显著更高效。在T1x的候选池偏移案例中,基于预训练MLIP嵌入和NTK的采集方法保持鲁棒,而委员会方法方差更高。结果表明,单个预训练MLIP即可实现可扩展、力感知且分布鲁棒的主动学习,用于基础模型微调。
Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs
Active learning for machine-learning interatomic potentials (MLIPs) must address several challenges to be practical: scaling to large candidate pools, leveraging energy-force supervision, and maintaining robustness when candidate pools are biased relative to the target distribution. In this work, we jointly address these challenges. We first introduce a linearly scaling acquisition framework based on chunked feature-space posterior-variance shortlisting. By avoiding materialisation of the candidate and train set kernels, this approach enables screening of ~200k structures within hours and applies broadly to acquisition strategies that score candidates based on molecular similarity metrics. We then extend the Neural Tangent Kernel (NTK) to a force-aware setting via mixed parameter-coordinate derivatives, yielding a force NTK and a joint energy-force NTK that provide natural similarity metrics for vector-field prediction. We demonstrate the effectiveness of the joint energy-force NTK on the OC20 dataset, where force-aware acquisition is crucial: it achieves the lowest energy and force MAE and RMSE across all metrics and distribution splits. Across T1x, PMechDB, and RGD benchmarks, our force NTK methods remain competitive with established baselines while being significantly more efficient than committee-based approaches. Under a controlled candidate-pool shift case study on T1x, acquisition based on pretrained MLIP embeddings and NTKs remains robust, whereas committee-based methods exhibit higher variance. Overall, these results show that a single pretrained MLIP can enable scalable, force-aware, and distribution-robust active learning for foundation-model fine-tuning.