Organizations: Department of Computing, Imperial College London, London, UK · School of Electronic Engineering and Computer Science, Queen Mary University of London, London, UK · Department of Computer Science, Stanford University, Stanford, CA, USA · Department of Integrative Biotechnology, Yonsei University, Incheon, Korea
Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.
Figures & tables
Figure 1: Overview of molecular representations and our HGR framework. a , Example molecule used throughout the schematic. b , Representative molecular representations, illustrated by a SMILES string, a pairwise atom–bond graph and higher-order boundary matrices such as B0,1 , B0,2 and B1,2 , which encode incidence relations. c , Higher-order lifting produces a combinatorial complex that encodes the molecule’s hierarchical higher-order topology, with its cells forming a contraction hierarchy. For clarity, each cell is annotated only with the right-hand-side graph GiR of its associated rule. d , Each production rule ri:GiL→GiR maps a left-hand-side interface graph that specifies a non-terminal site and its ordered anchor interface to the corresponding right-hand-side local graph. Distinct rules induced across the molecular library form the shared corpus P={ri∣i=1,…,∣P∣} . e , HGR serialises each complex as an ordered string r of reusable rule tokens and thereby provides a common input representation for diverse molecular-learning models, such as generative and foundation models. f , Starting from the start symbol S , successive rule applications replace selected non-terminal sites with their corresponding GiR and reconnect boundary bonds through the anchor interfaces, thereby reconstructing the combinatorial complex and its associated molecular graph.
Figure 2: Ring-system sparsity and grammar-rule statistics. a , Scale of experimentally accessible chemical collections (left axis) compared with ring-system diversity observed in practice (right axis). b , Growth of the grammar corpus on ZINC250k, quantified as the number of unique rules discovered as molecules are progressively added during grammar induction. c , Frequency distributions of grammar rules induced on ZINC250k.
Figure 3: Ring-system diversity across molecular benchmarks. a , Ring Diversity Index (RDI). b , Number of unique ring systems. c , Quantitative Ring Complexity Index (QRCI) distributions. Higher QRCI values indicate greater ring complexity and topological diversity. The light-blue region (QRCI = 1.8–5) denotes the range typically observed for approved drugs, balancing structural diversity and synthetic feasibility. Values above this range indicate more complex ring systems that may reduce drug-likeness and synthetic accessibility. d , Fractions of molecules containing each ring-size or ring-connectivity category, together with the unique-scaffold ratio.
GuacaMol
Method
Val. ↑
Unique. ↑
Nov. ↑
FCD score ↑
KL score ↑
DiGress [ 42 ]
85.2
100.0
99.9
68.0
92.9
DeFoG [ 43 ]
99.0
99.0
–
73.8
97.7
HoG-Diff [ 23 ]
99.1
100.0
97.0
78.3
96.9
SELFIES-VAE [ 14 ]
100.0
100.0
99.7
39.9
88.0
HGR-VAE
100.0
100.0
99.4
83.7
93.8
HGR-LDF
100.0
100.0
98.9
84.4
95.7
Table 1: Generation performance across molecular benchmarks.
Figure 4: Generation fidelity and chemical-space coverage of molecular generative models on RingDiv300k. a , b , Functional-group prevalence in the RingDiv300k test set and HGR-generated molecules for common functional groups ( a ) and less frequent or special motifs ( b ). Bars show motif prevalence; error bars denote s.d. across three generation seeds. c , d , Size-resolved Jensen–Shannon (JS) divergence between generated and test-set distributions for ring-size distributions ( c ) and functional-group composition ( d ), with molecules binned by heavy-atom count. Lower JS indicates closer agreement with the test distribution; grey dotted curves show the train-versus-test reference. e , Lines show size-resolved synthetic-accessibility (SA) scores; shaded bands denote s.d. across generation seeds. Lower SA indicates easier synthesis. f , Chemical-space coverage in Morgan-fingerprint Uniform Manifold Approximation and Projection (UMAP) space. Filled densities show generated molecules for each model (HGR variants, green; baselines, blue), and black contours show the test-set density.
Figure 5: Workflow and transfer performance of the HGR-based molecular foundation model. a , Overview of HGR-based pretraining and downstream transfer. A pretrained Transformer with rank-aware rotary position encoding (RoPE) and an HGR attention bias yields a molecular fingerprint, jointly optimised with atom-, substructure-, 3D- and global-level pretext tasks and transferred to downstream tasks via lightweight multilayer-perceptron (MLP) task heads. b , Normalised AUCs after ZINC2m pretraining under probing (left) and full fine-tuning (right) across seven MoleculeNet benchmarks. AUCs are min–max normalised per task across all methods; annotated values are the raw HGR-FM AUCs. c , Ablation of higher-order rule tokens and attention bias. Curves show validation loss during pretraining for the full HGR-FM model and variants without each component. Inset, mean probing AUC across seven downstream tasks. d , Comparison of training regimes: no pretraining, probing and full fine-tuning. e , Effect of pretraining data source, comparing RingDiv, ZINC2m and their combination. Left, mean probing AUC; right, number of downstream tasks for which each setting ranks first or second.
Symbol
Description
G=(V,E)
Molecular graph with vertex set V (atoms) and edge set E (bonds).
U∼GW
Adjacency between vertex subsets: ∃(u,w)∈E with u∈U and w∈W .
C=(V,X,rk)
Combinatorial complex with cell family X (atom and lifted motif cells) and rank function rk ; the lifting variant is written Cη , η∈{MIG,RSG} , and dropped when clear.
M
Motif vocabulary mined from the molecule library to score merge candidates in lifting.
Γ=(Σ,N,S,P)
Higher-order grammar with terminal set Σ , non-terminal set N , start symbol S∈N , and production-rule corpus P .
N∗∈N
A generic non-terminal symbol marking replaceable sites; Nk∗ denotes the site marked by N∗ and created at parsing step k .
Supplementary Table 1: Notation summary.
Supplementary Algorithm 1 MIG lifting
Supplementary Fig. 1: Illustration of the grammar induction process. Cells are successively contracted into non-terminal sites, shown as squares, until only the start symbol S remains. Applying the recorded rules in reverse order reconstructs the lifted complex and its associated molecular graph, whereas grammar-constrained derivation supports generation.
Supplementary Algorithm 2 ParseComplex : bottom-up parsing of a lifted complex into an HGR string
Supplementary Fig. 2: Workflow of HGR-based molecular generative models. Molecules are parsed into HGR strings, encoded by the grammar encoder and reconstructed by the grammar decoder. An optional latent diffusion module enables sampling in the learned latent space.
Variant
N
SRing
NRing
MRing
Spiro
Fused
Condensed
Bridged
RingDiv
1,183,434
157,542
1,160,912
69,061
104,198
522,461
582,747
113,423
RingDiv300k
299,819
42,703
290,985
33,591
32,026
147,500
164,171
37,660
Supplementary Table 2: Ring-topology composition of the released RingDiv variants. Each cell reports the number of molecules containing at least one ring of the indicated type.
Supplementary Table 3: Categorization of evaluation metrics for molecular generative models.
Supplementary Fig. 3: Functional groups and SMARTS patterns for molecular profiling. Representative structures and RDKit-compatible SMARTS queries used to annotate functional groups in RingDiv300k reference and generated molecules. Categories follow a previous functional-group diversity analysis and include aryl and halogen motifs, nitrogen-containing groups, oxygen- and carbonyl-containing groups, aromatic motifs, sulfur-containing groups and strained small rings. SMARTS patterns were adapted from prior functional-group profiling work [ 87 ] and applied as molecule-level, non-exclusive queries.
QM9
ZINC250k
RingDiv300k
MOSES
GuacaMol
Architecture
Max rule-sequence length
17
72
137
52
175
Production rules ∣P∣
280
403
6,205
320
2,385
Motif vocabulary
500
500
1,000
1,000
1,000
Latent dimension d
32
256
256
256
768
Rule-token dimension
64
256
768
512
768
Supplementary Table 4: Grammar variational autoencoder (HGR-VAE) configuration for the five generation benchmarks under the RSG-induced grammar.
Supplementary Table 6: HGR-FM pretraining and fine-tuning configuration.
Supplementary Fig. 4: Grammar statistics on QM9 dataset. a , Grammar corpus growth on QM9, showing the number of unique rules discovered as progressively more molecules are included during grammar induction. b , Rule-frequency distributions on QM9.
Model
Val. ↑
Unique. ↑
Nov. ↑
V.U.N.
KL score ↑
FCD ↓
NSPDK ↓
SA dist. ↓
κBF↓
κOR↓
Training
100.0 ± 0.0
100.0 ± 0.0
N/A
N/A
99.8 ± 0.0
0.1737 ± 0.0024
0.0001 ± 0.0000
0.0123 ± 0.0040
0.1312 ± 0.0577
0.1203 ± 0.0644
DiGress
98.9 ± 0.0
100.0 ± 0.0
100.0 ± 0.0
98.90 ± 0.04
98.2 ± 0.1
0.6300 ± 0.0110
0.00041 ± 0.00006
0.4171 ± 0.0247
0.1021 ± 0.0316
0.2656 ± 0.0582
DeFoG
98.3 ± 0.0
100.0 ± 0.0
100.0 ± 0.0
98.27 ± 0.04
98.7 ± 0.1
0.8575 ± 0.0087
0.00109 ± 0.00002
0.1298 ± 0.0081
0.4735 ± 0.0373
0.4045 ± 0.0088
HoG-Diff
100.0 ± 0.0
99.9 ± 0.0
99.6 ± 0.1
99.57 ± 0.1
98.2 ± 0.1
0.9250 ± 0.0236
0.00134 ± 0.00026
0.1435 ± 0.0333
0.1810 ± 0.0326
0.1729 ± 0.03247
SMILES-VAE
88.1 ± 0.1
100.0 ± 0.0
100.0 ± 0.0
88.1 ± 0.1
97.6 ± 0.0
0.6642 ± 0.0156
0.00072 ± 0.00005
0.1653 ± 0.0117
0.1556 ± 0.0061
0.2375 ± 0.0502
SELFIES-VAE
100.0 ± 0.0
100.0 ± 0.0
100.0 ± 0.0
100.0 ± 0.0
92.8 ± 0.4
2.4266 ± 0.0526
0.00286 ± 0.00005
0.6077 ± 0.0071
0.6891 ± 0.0687
0.3874 ± 0.0738
Supplementary Table 7: Generation performance on the RingDiv300k benchmark.
Model
Val. ↑
Unique. ↑
Nov. ↑
V.U.N.
FCD score ↑
KL score ↑
κBF↓
κOR↓
Training
100.0 ± 0.0
100.0 ± 0.0
N/A
N/A
96.12 ± 0.09
99.71 ± 0.00
0.1276 ± 0.0819
0.0573 ± 0.0304
DiGress
85.2
100.0
99.9
-
68.0
92.9
-
-
Disco
86.6
86.6
86.5
-
59.7
92.6
-
-
Cometh
98.9
98.9
97.6
-
72.7
96.7
-
-
DeFoG
99.0
99.0
-
97.9
73.8
97.7
-
-
HoG-Diff
99.1 ± 0.1
100.0 ± 0.0
97.0 ± 0.3
95.9 ± 0.2
78.3 ± 0.3
96.9 ± 0.3
0.1726 ± 0.0937
0.0944 ± 0.0423
Supplementary Table 8: Generation performance on the GuacaMol benchmark.
Method
Val. ↑
Unique. ↑
Nov. ↑
V.U.N. ↑
FCD ↓
NSPDK ↓
κBF↓
κOR↓
Training
100.00 ± 0.00
99.99 ± 0.00
N/A
N/A
0.074 ± 0.004
0.0002 ± 0.0000
0.1337 ± 0.0560
0.2380 ± 0.0762
MiCaM
99.93
93.89
83.25
–
1.045
0.001
–
–
MoFlow
91.36 ± 1.23
98.65 ± 0.57
94.72 ± 0.77
–
4.467 ± 0.595
0.017 ± 0.003
–
–
EDP-GNN
47.52 ± 3.60
99.25 ± 0.05
86.58 ± 1.85
–
2.680 ± 0.221
0.005 ± 0.001
–
–
GraphEBM
8.22 ± 2.24
97.90 ± 0.05
97.01 ± 0.17
–
6.143 ± 0.411
0.030 ± 0.004
–
–
GDSS
95.72 ± 1.94
98.46 ± 0.61
86.27 ± 2.29
86.48
2.900 ± 0.282
0.003 ± 0.000
0.9252
0.6012
Supplementary Table 9: Generation performance on the QM9 benchmark.
Method
Val. ↑
Unique. ↑
Nov. ↑
V.U.N. ↑
FCD ↓
NSPDK ↓
κBF↓
κOR↓
Training
100.00 ± 0.00
99.98 ± 0.01
N/A
N/A
0.201 ± 0.005
0.0001 ± 0.0000
0.1007 ± 0.0563
0.0478 ± 0.0190
MiCaM
100.00
88.48
99.98
–
32.395
0.166
–
–
MoFlow
63.11 ± 5.17
99.99 ± 0.01
100.00 ± 0.00
–
20.931 ± 0.184
0.046 ± 0.002
–
–
EDP-GNN
82.97 ± 2.73
99.79 ± 0.08
100.00 ± 0.00
–
16.737 ± 1.300
0.049 ± 0.006
–
–
GraphEBM
5.29 ± 3.83
98.79 ± 0.15
100.00 ± 0.00
–
35.471 ± 5.331
0.212 ± 0.075
–
–
GDSS
97.01 ± 0.77
99.64 ± 0.13
100.00 ± 0.00
–
14.656 ± 0.680
0.019 ± 0.001
–
–
Supplementary Table 10: Generation performance on the ZINC250k benchmark.
Method
Val. ↑
Unique. ↑
Nov. ↑
V.U.N. ↑
FCD ↓
Filters ↑
SNN ↑
Scaf. ↑
κBF↓
κOR↓
Training
100.0 ± 0.0
100.0 ± 0.0
-
-
0.518 ± 0.007
100.0 ± 0.1
0.586 ± 0.000
-
0.1138 ± 0.0303
0.1365 ± 0.01592
GraphINVENT
96.4
99.8
-
-
1.22
95.0
0.54
12.7
-
-
DiGress
85.7
100.0
95.0
-
1.19
97.1
0.52
14.8
-
-
DisCo
88.3
100.0
97.7
-
1.44
95.6
0.50
15.1
-
-
Cometh
90.5
99.9
92.6
-
1.27
99.1
0.54
16.0
-
-
DeFoG
92.8
99.9
92.1
-
1.95
98.9
0.55
14.4
-
-
Supplementary Table 11: Generation performance on the MOSES benchmark.
Supplementary Table 12: Transfer performance of pretrained molecular encoders.
Supplementary Fig. 5: Accelerated evaluation of generation metrics. a , Wall-clock execution time of representative evaluation metrics on the RingDiv300k benchmark (Scaf, SNN, Frag, FCD, NSPDK, κBF , κOR ). Dark bars indicate the optimised implementation; light shading denotes time saved relative to the original pipeline. Blue bold numbers above bars denote the speed-up factor, and grey numbers report the post-optimisation runtime. b , Runtime of the two slowest metrics ( κBF and κOR ) across additional standard datasets (QM9, ZINC250k, MOSES and GuacaMol), highlighting consistent acceleration under diverse data regimes. Data represent mean values for ten independent replicates.
Method
QM9
ZINC250k
MOSES
GuacaMol
RingDiv300k
Training time (h)
DiGress
-
-
-
-
63.0
DeFoG
-
-
-
-
53.1
HoG-Diff
3.1
6.5
7.8
30.0
8.1
HGR-VAE (RSG)
0.3
3.1
1.4
8.4
3.0
Inference time (min)
Supplementary Table 13: Training and inference runtime across molecular datasets.
Despite the success of foundation models in language and vision, molecular graph generation still lacks a unified framework for heterogeneous design tasks with reliable controllability. While reinforcement learning (RL) offers a natural post-training mechanism for task-specific optimization, applying it to graph generative models is hindered by the vast atom-wise action spaces and chemically invalid intermediate states. We propose \textbf{Co}ntrollable \textbf{Mole}cular Generative Foundation Models (CoMole), built with a unified motif-aware graph diffusion pipeline. By learning a motif-aware graph space, CoMole transfers pretrained structural priors into controllable generation, where RL optimizes conditional reverse policies over chemically meaningful decisions. We theoretically characterize the bottleneck of atom-level RL and justify motif-aware policy optimization. Across three heterogeneous benchmarks spanning materials and drug discovery, CoMole ranks first in controllability on all nine targets, reduces MAE by up to 48.2% relative to the strongest baselines, and maintains validity above 0.94 without rule-based correction or post-hoc filtering. We further show that CoMole transfers controllability to unseen properties by optimizing only task embeddings with the generator frozen, achieving performance competitive with strong task-specific baselines.
Transformer-based language models for SMILES strings suffer from a locality gap: standard character-level tokenization fragments chemically meaningful motifs, forcing models to repeatedly learn local syntax at the expense of long-range dependencies. To address this without disrupting standard tokenizers, we propose MolGram, which integrates a conditional n-gram memory module into molecular language models. MolGram maps local string patterns to learned embeddings via scalable hash lookups and dynamically injects this regional context into hidden states. Evaluations across three tasks, including unconditional molecule generation, forward reaction prediction, and single-step retrosynthesis, show that MolGram consistently improves performance. Crucially, our analyses demonstrate that MolGram outperforms baselines with 3× more parameters, establishing explicit local pattern memory as a highly efficient inductive bias.
Xinni Zhang, Zijing Liu, He Cao +2
1The Chinese University of Hong Kong · 2International Digital Economy Academy
Language models for molecular design have scaled to hundreds of millions of parameters, yet how they learn chemical grammar is poorly understood. We train SMolLM, a 53K-parameter weight-shared transformer, to generate novel SMILES with 95% validity on the ZINC-250K drug-like-molecule benchmark, outperforming a standard GPT with 10 times more parameters. Mechanistically, the same block resolves SMILES constraints across passes in a fixed hierarchy: brackets first, rings second, and valence last, as shown by error classification and linear probing, with ablation isolating the bracket-matching head. Together, these results yield a compact, mechanistically interpretable molecular generator and a testbed for studying iterative computation in formal-language domains.