Neural Petri Flow:用 Petri 网语义学习化学反应预测
Neural Petri flows for chemical reactions
化学+AI方向的朋友可以看看,NPF把Petri网直接做成网络结构,零样本原子映射就超过了RXNMapper,反应产物预测还保证分子合法。
Neural Petri Flow(NPF)把 Petri 网结构硬编码为无参数层,只训练变迁的速率律或分类读出。无需训练的零样本设定下,最小发射向量在 Golden set 上原子映射准确率 88.8%,高于 RXNMapper 的 85.6%;在 EnzymeMap 酶反应上是 88.7% 对 77.9%。在 USPTO-480K 上训练后产物预测达 87.7%,仅用 1% 训练子集也有 67.4%。ECREACT 的 EC 编号三级预测准确率 90.2%,比已发表最好方法高 5.6 个百分点。用电子作 token 时对 FlowER 基本步骤的 top-1 预测 90.5%,且每个预测都是无需过滤的合法分子。
Neural Petri flows for chemical reactions
Petri nets have been used to describe chemical processes such as reactions.They map well to chemistry: Places are the bonds between atoms and the free valence of each atom, a token is a unit of bond order, a transition forms or breaks a bond, the conserved quantities are the valence budgets of the atoms, and the enabling rule is the valence rule. These semantics are not guaranteed by learned models of reactions or neural networks that are built on Petri nets that use the net as a scaffold for message passing. Here, we ask what architecture remains a Petri net for every value of its weights. We find the answer in the theory, where all semantics of a net share the firing form $m^\prime=m+Cσ$, locality, as enabling reads only the inputs of a transition, and the enabling rule, and we prove that conservation forces the firing form and that non-negativity forces the enabling rule on local rate laws. This leaves free the rate law, which is the propensity of each transition to fire. We introduce Neural Petri Flow, which learns this rate law, or a readout for classification, and hard-wires the rest as parameter-free layers. On what we denote a valence net, atom mapping, reaction classification, and forward prediction become three tasks on one firing vector. Without training, the minimum firing vector maps 88.8% of the curated Golden set against 85.6% for RXNMapper, and 88.7 against 77.9% of the enzymatic reactions of EnzymeMap. On USPTO-480K, NPF trained on these firing vectors predicts 87.7% of the products and 67.4% when trained on a 1% subset of the training reactions. EC numbers of ECREACT are predicted at the third level for 90.2% of reactions, 5.6 points ahead of the best published method. With electrons as tokens, the same token game predicts 90.5% of the elementary steps of FlowER first, ahead of the published baseline, and every top-1 prediction is a valid molecule without a filter.