APO:无监督原子策略优化,用于原子系统 3D 结构预测

APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

精选理由

想搞材料或药物结构预测的可以看看:APO 不用实验标签,纯无监督也能超过全监督的 FlowDPO。

AI 摘要

论文提出无监督对齐框架 APO,用于预测原子系统 3D 结构,不需要 ground-truth 参考结构。APO 将 group-relative policy optimization 适配到 3D 原子环境,用双奖励机制强化潜在结构模式和热力学稳定性。在晶体和抗体结构预测基准上,APO 持续优于全监督基线 FlowDPO,达到新的最优匹配率和结构保真度。APO 还通过拉直概率路径提升了推理效率。

原文 · arXiv cs.LG

APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) have recently shown promise in this domain, their performance relies heavily on alignment with ground-truth coordinates via supervised preference learning. However, obtaining experimental labels for novel crystal phases or de novo proteins is prohibitively expensive, creating a bottleneck for structural modeling in data-scarce regimes. In this work, we propose (Atomic Policy Optimization), a fully unsupervised alignment framework that eliminates the need for ground-truth reference structures. APO adapts group-relative policy optimization to 3D atomic environments, utilizing a novel dual-reward mechanism: (i) a that reinforces the policy's dominant latent structural modes through eigen-decomposition of sample similarities, and (ii) a that enforces thermodynamic stability. Our framework enables the model to ``self-correct'' by identifying physically plausible configurations within sampled groups. Extensive benchmarks on crystal and antibody structure prediction demonstrate that APO consistently outperforms fully supervised baselines, achieving a new state-of-the-art in match rates and structural fidelity. Furthermore, we show that APO effectively straightens probability paths, significantly improving inference efficiency. Our results suggest that intrinsic physical consistency can serve as a superior guide for alignment compared to noisy, supervised coordinate matching.