Neural Petri flows for chemical reactions
Organizations: Bioinformatics Group Bioprocess Engineering Group Wageningen Unversity Wageningen, the Netherlands
Abstract
Petri nets have been used to describe chemical processes such as reactions.They map well to chemistry: Places are the bonds between atoms and the free valence of each atom, a token is a unit of bond order, a transition forms or breaks a bond, the conserved quantities are the valence budgets of the atoms, and the enabling rule is the valence rule. These semantics are not guaranteed by learned models of reactions or neural networks that are built on Petri nets that use the net as a scaffold for message passing. Here, we ask what architecture remains a Petri net for every value of its weights. We find the answer in the theory, where all semantics of a net share the firing form , locality, as enabling reads only the inputs of a transition, and the enabling rule, and we prove that conservation forces the firing form and that non-negativity forces the enabling rule on local rate laws. This leaves free the rate law, which is the propensity of each transition to fire. We introduce Neural Petri Flow, which learns this rate law, or a readout for classification, and hard-wires the rest as parameter-free layers. On what we denote a valence net, atom mapping, reaction classification, and forward prediction become three tasks on one firing vector. Without training, the minimum firing vector maps 88.8% of the curated Golden set against 85.6% for RXNMapper, and 88.7 against 77.9% of the enzymatic reactions of EnzymeMap. On USPTO-480K, NPF trained on these firing vectors predicts 87.7% of the products and 67.4% when trained on a 1% subset of the training reactions. EC numbers of ECREACT are predicted at the third level for 90.2% of reactions, 5.6 points ahead of the best published method. With electrons as tokens, the same token game predicts 90.5% of the elementary steps of FlowER first, ahead of the published baseline, and every top-1 prediction is a valid molecule without a filter.
Figures & tables
| method | learned | Golden | NatComm | USPTO-3k | Recon3D | E. coli |
|---|---|---|---|---|---|---|
| RDT | rules | 82.54 | 84.11 | 90.87 | 54.97 | 78.02 |
| RXNMapper | yes | 87.43 | 87.58 | 93.53 | 48.69 | 72.53 |
| LocalMapper | yes | 89.08 | 92.67 | 97.77 | 50.79 | 69.96 |
| GraphormerMapper | yes | 89.59 | 92.87 | 95.10 | 34.82 | 42.12 |
| NPF, minimum firing vector | no | 89.70 | 90.63 | 92.13 | 57.85 | 72.89 |
| Schneider50k (50 classes) accuracy macro-F1 DRFP + MLP 3 PGNN, readout 3 PGNN, firing vector 5 NPF, readout 5 NPF, firing vector 5 | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| ECREACT, Enzyformer split | EC4 | EC3 | EC2 | EC1 | CARE task 2, easy | EC4 | EC3 | EC2 | EC1 |
| BEC-Pred | – | 76.9 | 81.2 | 85.8 | CLIPZyme | 12.2 | 39.9 | 61.8 | 79.9 |
| DRFP + CLM | – | 81.7 | 85.6 | 88.3 | DRFP similarity | 59.3 | 77.1 | 85.2 | 90.6 |
| Synthcoder-DistilBERT | – | 84.4 | 86.7 | 90.9 | CREEP | 39.4 | 66.4 | 79.9 | 92.9 |
| Enzyformer | – | 84.6 | 88.2 | 90.9 | NPF, no map | ||||
| NPF | NPF + mapper | ||||||||
| method | prediction modality | top-1 | top-3 | top-5 |
|---|---|---|---|---|
| Molecular Transformer † ( Schwaller et al., 2019 ) | SMILES | 0.886 | 0.935 | 0.942 |
| Graph2SMILES † ( Tu and Coley, 2022 ) | SMILES | 0.903 | 0.940 | 0.948 |
| MEGAN † ( Sacha et al., 2021 ) | Graph edits | 0.863 | 0.924 | 0.940 |
| GTPN † ( Do et al., 2019 ) | Graph edits | 0.832 | 0.860 | 0.865 |
| NERF † ( Bi et al., 2021 ) | Electron edits | 0.907 | 0.933 | 0.937 |
| MAELLE ( Xuan-Vu et al., 2026 ) | Electron flow | 0.872 | 0.930 | 0.939 |
| step, top- | pathway, top- | ||||||||
| method | params | 1 | 2 | 3 | 5 | 1 | 2 | 3 | 5 |
| Molecular Transformer | 12M | 88.75 | 96.62 | 98.42 | 98.92 | 88.31 | 95.64 | 97.02 | 97.59 |
| Graph2SMILES | 18M | 89.09 | 96.75 | 98.08 | 98.66 | 92.51 | 95.94 | 97.03 | 97.76 |
| Graph2SMILES+H | 18M | 87.39 | 95.31 | 96.87 | 97.67 | 89.22 | 93.70 | 95.10 | 96.19 |
| FlowER | 7M | 88.48 | 96.42 | 97.96 | 98.60 | 88.97 | 94.75 | 96.48 | 97.39 |
| FlowER-large | 16M | 89.74 | 97.40 | 98.66 | 99.13 | 92.50 | 96.89 | 98.15 | 98.61 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| method | within studies | new study | new carbon source | balance |
|---|---|---|---|---|
| NPF, GMA rate law | 0.114 0.002 | 0.153 0.013 | 0.385 0.002 | 0.006 |
| NPF, rate law described in Section C.1 | 0.149 0.007 | 0.157 0.019 | 0.363 0.023 | 0.019 |
| PGNN, GMA rate law | 0.119 0.008 | 0.260 0.019 | 0.798 0.053 | 2 216 |
| PGNN | 0.144 0.023 | 0.280 0.019 | 0.858 0.070 | 1 009 |
| PGNN+ with the state equation, GMA | 0.211 0.017 | 0.208 0.072 | 0.399 0.025 | 0.008 |
| random forest, projected onto the net | 0.095 | 0.161 | 0.350 | 0.000 |
| method | within studies | new study | new carbon source | balance |
|---|---|---|---|---|
| NPF, GMA rate law | 0.170 0.013 | 0.245 0.085 | 0.387 0.020 | 0.011 |
| PGNN, GMA rate law | 0.120 0.003 | 0.256 0.020 | 0.979 0.438 | 2 052 |
| PGNN+ with the state equation, GMA | 0.265 0.088 | 0.355 0.141 | 0.348 0.016 | 2.512 |
| random forest, projected onto the net | 0.117 | 0.154 | 0.333 | 0.000 |
| random forest | 0.112 | 0.175 | 0.869 | 1.314 |
| ridge regression, projected onto the net | 0.144 | 0.176 | 0.649 | 0.000 |
| cost | choice | held out | all | interval | proved |
|---|---|---|---|---|---|
| bonds made and broken only | 83.0 | 81.03 | 81.25 | 79.4 to 83.1 | 99.94 |
| + second level, all hydrogen | 83.5 | 82.44 | 82.56 | 80.8 to 84.3 | 99.89 |
| + hydrogen on heteroatoms not counted | 85.0 | 83.40 | 83.58 | 81.8 to 85.3 | 99.89 |
| + C–H as bond places (the cost) | 86.0 | 83.65 | 83.92 | 82.2 to 85.6 | 100.00 |
| + third level (the mapper) | 91.0 | 88.53 | 88.81 | 87.3 to 90.2 | 100.00 |
| RXNMapper | 85.5 | 85.58 | 85.57 | 83.9 to 87.2 | – |
| places | reactions | NPF | RXNMapper | only one | not minimal | |
|---|---|---|---|---|---|---|
| 1 or 2 | 772 | 95.1 | 92.7 | 38 / 20 | 0.025 | 0.4 |
| 3 or 4 | 658 | 87.8 | 84.8 | 46 / 26 | 0.024 | 3.6 |
| 5 or 6 | 212 | 82.1 | 74.5 | 30 / 14 | 0.023 | 3.8 |
| 7 or more | 118 | 65.3 | 62.7 | 16 / 13 | 0.711 | 12.7 |
| set | RDT | RXNMapper | LocalMapper | GraphormerMapper |
|---|---|---|---|---|
| Golden | 175 / 49, 0.001 | 109 / 69, 0.003 | 119 / 108, 0.51 | 122 / 120, 0.95 |
| NatComm | 41 / 9, 0.001 | 24 / 9, 0.01 | 21 / 31, 0.21 | 21 / 32, 0.17 |
| USPTO-3k | 78 / 40, 0.001 | 3 / 45, 0.001 | 4 / 173, 0.001 | 17 / 106, 0.001 |
| Recon3D | 15 / 4, 0.02 | 35 / 0, 0.001 | 30 / 3, 0.001 | 90 / 2, 0.001 |
| E. coli | 17 / 31, 0.06 | 30 / 29, 1.00 | 22 / 14, 0.24 | 90 / 6, 0.001 |
| criterion | NPF, minimum firing vector | RXNMapper | only one | |
| reaction centre (SynRXN) | 88.70 | 77.92 | / |
| method | accuracy | macro-F1 | |
|---|---|---|---|
| DRFP + MLP | 3 | ||
| generic readout | 3 | ||
| generic head, firing vector | 5 | ||
| NPF, state-equation readout | 5 | ||
| NPF, firing vector of the mapper | 5 |
| 1 000 labelled | 250 labelled | |||||
| model | full training set | larger molecules | alone | + firing task | alone | + firing task |
| DRFP + MLP | 95.52 0.04 | 78.43 0.63 | 86.62 0.77 | – | 65.03 1.58 | – |
| generic readout | 97.74 0.01 | 71.76 1.08 | 78.59 1.27 | 90.02 0.16 | 32.14 1.64 | 52.64 1.33 |
| NPF, state-equation readout without gate | 97.47 0.14 | 67.43 1.03 | 91.26 0.27 | 94.17 0.32 | 61.19 4.24 | 81.65 3.21 |
| NPF, state-equation readout | 97.68 0.09 | 71.55 3.65 | 91.57 0.64 | 94.76 0.20 | 57.01 3.19 | 86.77 1.26 |
| generic head on the firing vector | 98.64 0.09 | 82.83 1.03 | 93.83 0.51 | – | 48.99 1.31 | – |
| atom maps | accuracy | macro-F1 | (accuracy) | (macro-F1) |
|---|---|---|---|---|
| none, state-equation readout | 90.36 0.48 | 66.80 0.77 | – | – |
| RXNMapper | 90.47 1.00 | 67.23 1.96 | 0.87 | 0.75 |
| minimum firing vector (the mapper) | 91.04 0.74 | 67.51 1.54 | 0.26 | 0.53 |
| EnzymeMap’s curated maps | 91.81 0.12 | 69.63 0.58 | 0.03 | 0.008 |
| 1 % | 10 % | 25 % | ||||
|---|---|---|---|---|---|---|
| model | top-1 | top-5 | top-1 | top-5 | top-1 | top-5 |
| MT | 0.16 | 0.37 | 74.16 | 84.19 | 82.39 | 90.37 |
| NPF, supplied maps | 67.24 0.24 | 80.36 0.36 | 78.85 0.04 | 88.73 0.08 | – | – |
| NPF, targets of the mapper | 67.44 0.19 | 80.45 0.18 | 78.83 0.14 | 88.78 0.16 | 81.73 0.13 | 90.64 0.06 |
| NPF width 256, supplied maps | – | – | – | – | 82.95 0.26 | 91.04 0.12 |
| NPF width 256, targets of the mapper | – | – | – | – | 83.23 0.20 | 91.00 0.07 |
| targets | training reactions | greedy top-1 | beam top-1 |
|---|---|---|---|
| atom maps supplied in the data set | 78.57 0.04 | 79.05 0.04 | |
| mapper, set of tied vectors | 79.72 0.03 | 80.08 0.05 | |
| mapper, one vector | 79.15 0.26 | 79.54 0.25 |
| greedy top-1 | beam | |||||
|---|---|---|---|---|---|---|
| model | top-1 | top-5 | valid | |||
| NPF token game | 84.31 0.71 | 90.34 0.34 | 93.34 0.28 | 93.39 0.18 | 98.08 0.07 | 100.00 0.00 |
| without the enabling mask | 84.19 0.18 | 89.82 0.64 | 93.06 0.11 | 93.13 0.12 | 98.07 0.04 | 99.98 0.02 |
| one-shot, greedy valence repair | – | – | 81.60 0.50 | – | – | 100.00 0.00 |
| one-shot, generic | 66.28 0.51 | 75.02 0.53 | 79.72 0.22 | – | – | 97.21 0.30 |
| NPF model | valence rule | valid molecules | invalid | dropped fragments |
|---|---|---|---|---|
| aromatic, recorded maps | 100 | 99.875 | 50 | – |
| aromatic, targets of the net | 100 | 99.897 | 41 | 0.130 |
| aromatic, targets of the net, no enabling | 99.987 | 99.890 | 44 | 0.138 |
| method | valid SMILES | heavy atoms | atoms, protons, electrons |
|---|---|---|---|
| Molecular Transformer | 70.2 | 39.1 | 33.0 |
| Graph2SMILES | 76.3 | 30.7 | 17.2 |
| Graph2SMILES+H | 78.8 | 27.7 | 19.0 |
| FlowER | 95 | by construction | by construction |
| NPF, arrow net | 100 | by construction | by construction |
| NPF, electron net | 100 | by construction | by construction |
| component | criterion | Petri | generic | |
|---|---|---|---|---|
| one bond change at a time | top-1, all training reactions | 93.34 0.28 | 79.72 0.22 | 0.001 |
| same | top-1, 1 995 training reactions | 84.31 0.71 | 66.28 0.51 | 0.001 |
| enabling (Prop. 4 ) | valence-valid products | 100 (all weights) | 97.21 0.30 | – |
| enabling mask in the game | top-1 | 93.34 0.28 | 93.06 0.11 | 0.22 |
| readout (Prop. 3 ) | % changed by a far methyl | 0.00 | 0.27 | – |
| same | accuracy, 1 000 / 250 training reactions | 91.26 / 61.19 | 78.59 / 32.14 | 0.002 / 0.003 |