Authors: Chuqin Geng, Li Zhang, Mark Zhang, Zhaoyue Wang, Haolin Ye, Xujie Si
Organizations: School of Computer Science, McGill University, Montreal, Canada · Department of Computer Science, University of Toronto, Toronto, Canada
While deep generative models excel at capturing graph data distributions, they struggle to satisfy complex, hard constraints. In unconstrained settings, these models typically produce valid topologies; yet imposing strict compositional rules, like those in drug discovery, creates an out-of-distribution (OOD) setting where purely neural methods frequently fail. Because these neural approaches rely on soft conditioning and post-hoc filtering on such tasks, they cannot provide the formal guarantees needed for high-stakes domains. To address this, we introduce Neuro-Symbolic Graph Generative Modeling (NSGGM), a framework built on the principle of Neural Proposals, Symbolic Guarantees. NSGGM decouples generation: an autoregressive model proposes structural scaffolds, and a Satisfiability Modulo Theories (SMT) solver handles the final discrete assembly of the proposed substructures. Empirically, NSGGM is competitive with state-of-the-art methods on unconstrained tasks. To evaluate logical-constraint satisfaction inspired by real drug discovery workflows, we introduce MolSAT, a benchmark for hard compositional rules. On MolSAT, purely neural baselines completely fail OOD (0% satisfaction with zero training support), while NSGGM achieves >95% satisfaction in-distribution and 64-86% with zero training support.
Figures & tables
Figure 1 : An overview of the NSGGM framework.
Model
Valid ↑
Unique ↑
Novel ↑
FCD ↑
KL ↑
MoLeR [ 21 ]
100.0
100.0
99.1
62.5
96.4
MiCaM [ 8 ]
100.0
99.4
98.6
73.1
98.9
DiGress [ 29 ]
85.2
100.0
99.9
68.0
92.9
LSTM
95.9
100.0
91.2
91.3
99.1
DISCO-GT [ 32 ]
100.0
86.6
99.9
59.7
92.6
NSGGM (ours)
100.0±0.0
100.0±0.0
98.9±0.0
86.1±0.4
95.5±0.1
Table 1 : Unconditional de novo molecule generation on GuacaMol. FCD and KL are GuacaMol-normalized scores in [0,100] , higher is better.
Model
Valid ↑
Unique ↑
Novel ↑
Filters ↑
FCD ↑
SNN ↑
Scaf ↑
JT-VAE [ 12 ]
100.0
100.0
99.9
97.8
81.8
0.53
10.0
MoLeR [ 21 ]
100.0
99.9
97.1
—
85.2
—
—
DiGress [ 29 ]
85.7
100.0
95.0
97.1
78.8
0.52
14.8
GraphINVENT [ 23 ]
96.4
99.8
—
95.0
78.3
0.54
12.7
DISCO-GT [ 32 ]
88.3
100.0
97.7
95.6
74.9
0.50
15.1
NSGGM (ours)
100.0±0.0
100.0±0.0
94.4±0.3
97.1±0.1
82.3±0.3
0.5±0.0
11.2±0.1
Table 2 : Unconditional de novo molecule generation on MOSES ( TestSF split). Filters: fraction passing MOSES chemical filters. SNN: Tanimoto similarity to nearest test molecule. Scaf: Bemis–Murcko scaffold overlap. FCD: GuacaMol-normalized in [0,100] , higher is better.
Model
Valid ↑
Unique ↑
FCD ↑
KL ↑
JT-VAE [ 12 ]
100.0
54.9
58.8
89.1
MoLeR [ 21 ]
100.0
94.0
93.1
96.9
LO-ARM-st-sep [ 30 ]
99.9
98.9
95.3
—
DiGress [ 29 ]
99.0±0.1
96.2±0.1
—
—
UniGEM [ 7 ]
95.0±3.1
98.1
—
—
MiCaM [ 8 ]
100.0
93.2
94.5
98.0
Table 3 : Unconditioned molecule generation on QM9 implicit hydrogens. — indicates that the metric is not reported in the original paper.
Figure 2 : Two sample MolSAT instances solved by NSGGM . Left: a low-support constraint φ2 from Family III combining XOR, IFF, and IMP over substructure predicates. Right: a zero-support constraint from Family IV. In each panel, the highlighted substructures in the generated molecule correspond to the predicates referenced in the formula below, and the SMT solver returns a SAT certificate verifying the assignment.
I
II
III
IV
Method
H
L
Z
H
L
Z
H
L φ1
L φ2
Z
H
L
Z
MoLeR
14.1
7.1
0.0
23.0
4.3
0.0
13.3
10.9
5.9
0.0
12.5
14.2
0.0
GenMol
64.4
37.6
0.0
51.0
48.6
0.0
59.2
31.5
28.3
0.0
61.5
48.5
0.0
NSGGM (Ours)
98.9
99.9
79.7
100.0
99.9
84.1
96.7
99.7
96.0
63.9
99.6
98.9
85.6
Table 4 : Constraint satisfaction rate (Hit%) across 13 benchmarks in 4 families (I–IV), with tiers H = High, L = Low, and Z = Zero based on training coverage. MoLeR and GenMol receive a 106 -attempt budget per benchmark. Both neural baselines produce zero satisfying samples on all Zero-tier benchmarks. Best per column in bold.
#Motifs (k)
#Mols
#MergeVars (v)
#HardClauses
Neural (ms)
Solver med (ms)
Solver 95p (ms)
SAT Rate
2–5
519
16.5
26.3
52.6
3.1
5.2
1.000
6–10
6,294
62.3
98.0
79.9
10.5
28.5
1.000
11–15
2,832
150.7
242.9
113.3
40.7
107.6
1.000
16–20
322
332.1
644.1
154.3
154.2
367.5
0.997
21+
35
665.7
1,744.6
202.1
546.4
830.9
1.000
Overall
10,004
95.7
158.6
90.7
15.1
99.2
1.000
Table 5 : NSGGM solver scaling on unconstrained GuacaMol generation (10K molecules, single-core CPU), binned by blueprint size. #Motifs (k) = binned number of motif proposals per molecule; #MergeVars (v) = average candidate slots per molecule; #HardClauses = logical formula size; Neural = neural inference time per molecule; Solver med = median solve time.
Variant
FCD ↑
KL ↑
Frag. % ↓
Tokens Lost % ↓
Full
86.1
95.5
0.0
0.0
W/o parent pred.
49.2
87.6
62.5
23.4
W/o merge-type pred.
63.0
88.3
22.8
8.4
W/o blueprint char.
62.4
88.9
0.0
0.0
Table 6 : Per-component neural guidance ablation on GuacaMol ( 10,000 molecules; connectivity relaxed to measure fragmentation, with the remaining constraints held fixed across variants. Each row removes one guidance signal. W/o blueprint char. falls back to selecting the most frequent blueprint variant for each motif type. Frag. % : proportion of assembled graphs with disconnected components. Tokens Lost % : proportion of proposed tokens not merged into the largest connected component.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3 : Example outputs from QM9 implicit (left), MOSES (middle), and GuacaMol (right). Samples were randomly chosen.
I
II
III
IV
H
L
Z
H
L
Z
H
L φ1
L φ2
Z
H
L
Z
Train hits
12,370
1,276
0
7,215
2,448
0
26,704
605
387
0
17,275
5,784
0
Train freq (%)
0.972
0.100
0.000
0.567
0.192
0.000
2.098
0.048
0.030
0.000
1.357
0.454
0.000
Appendix
Table 7 : Natural prevalence of each logical constraint in the GuacaMol training set.
Symbol
Name
SMILES
Description
Pyr
Pyridine
C1=CC=NC=C1
Six-membered ring, one N
Pym
Pyrimidine
C1=CN=CN=C1
Six-membered ring, two N (1,3)
Imi
Imidazole
C1=CN=CN1
Five-membered ring, two N (1,3)
Thz
Thiazole
C1=CSC=N1
Five-membered ring, N + S
Qui
Quinoline
C1=CC=C2C(=C1)C=CC=N2
Fused benzene + pyridine
BzI
Benzimidazole
C1=CC=C2C(=C1)NC=N2
Fused benzene + imidazole
Appendix
Table 8 : Core scaffold predicates. Each predicate Pi corresponds to a Boolean atom xi that is true iff the candidate molecule contains the SMILES substructure as a subgraph match.
Symbol
Name
SMILES
Description
Ntr
Nitrile
C#N
Cyano group
Nto
Nitro
N+[O-]
Nitro group
TF3
Trifluoromethyl
C(F)(F)F
CF 3 group
Sul
Sulfonamide
S(=O)(=O)N
Sulfonamide group
Appendix
Table 9 : Functional motif predicates. Compact substituent groups used as electron-withdrawing or polar functional handles.
I
II
III
IV
Method
H
L
Z
H
L
Z
H
L φ1
L φ2
Z
H
L
Z
NSGGM (Ours)
0.769
0.803
4.605
0.536
1.006
5.468
4.898
0.585
0.969
5.907
0.666
0.951
8.933
MoLeR
0.895
1.446
∞
0.506
2.571
∞
0.890
1.167
2.259
∞
0.970
0.892
∞
GenMol
0.039
0.068
∞
0.047
0.054
∞
0.039
0.078
0.090
∞
0.038
0.052
∞
Appendix
Table 11 : Timing table for each model on MolSat. Reported time is number of seconds / valid molecule.
Model
Deg ↓
Clus ↓
Orb ↓
V.U.N. ↑
Stochastic block model
GraphRNN
6.9
1.7
3.1
5%
GRAN
14.1
1.7
2.1
25%
GG-GAN
4.4
2.1
2.3
25%
SPECTRE
1.9
1.6
1.6
53%
ConGress
34.1
3.1
4.5
0%
Appendix
Table 12 : Unconditional generation on SBM and planar graphs. V.U.N.: valid, unique & novel graphs.
Generating realistic and diverse graphs is a key problem in machine learning, with applications in molecular discovery, circuit design, cybersecurity, and beyond. However, current graph generative models remain limited by scalability and novelty. Diffusion-based methods often require costly full-adjacency operations and long denoising chains, while many autoregressive and hybrid models have at least quadratic complexity. In addition, these models often imitate training graphs rather than generalize beyond them. We propose a lightweight autoregressive framework to address these issues. It uses a structure-guided topological ordering to serialize graphs into regular edge sequences, enabling near log-linear generation, and a two-phase training strategy that combines exploration-oriented augmentation with iterative refinement to reduce overfitting and promote controlled novelty. Experiments on molecular and non-molecular benchmarks show that our approach improves novelty while preserving high validity and uniqueness. The framework also supports both LSTM and Mamba-style causal sequence backbones, with large-memory accelerators enabling longer graph-sequence experiments beyond typical GPU limits.
Neural Markov Logic Networks (NMLNs) are a flexible neurosymbolic relational model. Previous work has shown that, although NMLNs achieve strong performance as generative models for small relational structures, they underperform diffusion-based generative graph models on larger structures. In this paper, we strengthen NMLNs along two main dimensions: (i) we increase the expressive capacity of their potential functions using graph neural networks, and (ii) we develop a new training and inference algorithm inspired by parallel-tempering Markov chain Monte Carlo methods, which we name parallel noising. Together, these enhancements enable NMLNs to attain strong performance in graph generation relative to general diffusion-based generative graph models. Furthermore, they allow NMLNs to match the performance of specialized text-based recurrent models when generating small molecular structures.
Peter Jung, Giuseppe Marra, Ondrej Kuzelka
Czech Technical University, Prague, Czech Republic · Department of Computer Science, KU Leuven, Leuven, Belgium
Despite the success of foundation models in language and vision, molecular graph generation still lacks a unified framework for heterogeneous design tasks with reliable controllability. While reinforcement learning (RL) offers a natural post-training mechanism for task-specific optimization, applying it to graph generative models is hindered by the vast atom-wise action spaces and chemically invalid intermediate states. We propose \textbf{Co}ntrollable \textbf{Mole}cular Generative Foundation Models (CoMole), built with a unified motif-aware graph diffusion pipeline. By learning a motif-aware graph space, CoMole transfers pretrained structural priors into controllable generation, where RL optimizes conditional reverse policies over chemically meaningful decisions. We theoretically characterize the bottleneck of atom-level RL and justify motif-aware policy optimization. Across three heterogeneous benchmarks spanning materials and drug discovery, CoMole ranks first in controllability on all nine targets, reduces MAE by up to 48.2% relative to the strongest baselines, and maintains validity above 0.94 without rule-based correction or post-hoc filtering. We further show that CoMole transfers controllability to unseen properties by optimizing only task embeddings with the generator frozen, achieving performance competitive with strong task-specific baselines.