Latent Flow Matching for Molecular Graph Generation
Organizations: Mohamed bin Zayed University of Artificial Intelligence
Abstract
Modern graph generative models typically operate directly in the discrete graph space, explicitly generating node and edge variables, which can become costly as graphs grow. In this paper, we perform generation explicitly on latent representations of entire graphs obtained from a pretrained Variational Autoencoder with high reconstruction fidelity. The generated representations, obtained through flow matching, are then decoded only at the final step. Across molecular benchmarks of increasing size, our approach achieves strong validity and FCD while offering a favorable quality-efficiency trade-off compared with state-of-the-art explicit graph generative models. One of the main advantages of this formulation is that the graph representation only needs to be learned once, after which the same one can be reused across multiple generative objectives without retraining. We demonstrate generation guided by molecular properties and further introduce validity-aware generation though a classifier learned directly in latent space. All code will be made available upon acceptance.
Figures & tables
| Name | # Graph | # atom | # bonds | Max Size | Atom types | Emb. Dim. |
| QM9 | k | |||||
| PubChem16 | M | |||||
| PubChem32 | M |
| Method | Sampler | NFE | Validity | Uniq. | Novelty | FCD | Time (s) | VUN/T |
| QM9 | ||||||||
| DiGress | Diffusion | 2 | N/A | |||||
| DiGress | Diffusion | 5 | N/A | |||||
| DiGress | Diffusion | 20 | N/A | |||||
| DiGress | Diffusion | 500 | N/A | 0.061 | ||||
| CFM-CSD | CSD | 1 | N/A |
| Method / component | QM9 | PubChem16 |
| DiGress | 7 h 34 min | 16 h 55 min |
| CFM-CSD | 63 h 49 min | 73 h 21 min |
| CFM-ECLD | 70 h 53 min | 71 h 40 min |
| Latent FM | 27 min | 3 h 50 min |
| Validity dataset construction | 2 h 06 min | 8 h 05 min |
| Validity predictor | 1 min 28 s | 1 min 12 s |
| Dataset | Property | Conditional FM | + Guidance | ||
| MAE | MAE | ||||
| QM9 | logP | 0.100 | 0.978 | 0.063 | 0.988 |
| QM9 | MW | 1.705 | 0.979 | 1.674 | 0.981 |
| PubChem16 | logP | 0.282 | 0.966 | 0.154 | 0.978 |
| PubChem16 | MW | 3.040 | 0.976 | 2.578 | 0.977 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Steps | Validity | Uniqueness | Novelty | FCD |
| CFM-CSD | 1 | ||||
| CFM-CSD | 2 | ||||
| CFM-CSD | 5 | ||||
| CFM-CSD | 20 | ||||
| CFM-ECLD | 1 | ||||
| CFM-ECLD | 2 |
| Dataset | Property | Validity | Uniqueness |
| QM9 | logP | ||
| QM9 | MW | ||
| PubChem16 | logP | ||
| PubChem16 | MW |
| Representation | Property | MAE | |
| AE | logP | 0.247 | 0.966 |
| VAE | logP | 0.154 | 0.978 |
| AE | MW | 3.923 | 0.971 |
| VAE | MW | 2.578 | 0.977 |
| Hyperparameter | QM9NoHydro | PubChem16 |
| Latent dimension | 64 | 256 |
| Latent transform (FM) | None | Standardization |
| Flow Matching | ||
| Hidden dimension | 1024 | 2048 |
| Number of layers | 6 | 10 |
| Time embedding dimension | 1 | 64 |
| Hyperparameter | QM9 MW | QM9 logP |
| Latent dimension | 64 | 64 |
| Conditioning property | MW | logP |
| Property normalization | Standardized | Standardized |
| Latent transform | None | None |
| Conditional Flow Matching | ||
| Hidden dimension | 512 | 512 |
| Hyperparameter | PubChem16 MW | PubChem16 logP |
| Latent dimension | 256 | 256 |
| Conditioning property | MW | logP |
| Property normalization | Standardized | Standardized |
| Latent transform | None | None |
| Conditional Flow Matching | ||
| Hidden dimension | 1024 | 1024 |