The increasing variety of molecular data creates a need for generative models that can flexibly incorporate heterogeneous constraints across modalities. However, existing SMILES-based diffusion models are typically designed for a fixed conditioning modality, and introducing new constraints often requires retraining the model. To address this limitation, we propose Cross-Modality Controlled Molecule Generation with Diffusion Language Model (CMCM-DLM), a modular framework that extends a pre-trained diffusion model to support heterogeneous molecular constraints without retraining the backbone. We demonstrate CMCM-DLM using two complementary modalities: molecular structure and chemical properties. Specifically, a Structure Control Module (SCM) guides early diffusion steps to establish the molecular scaffold, while a Property Control Module (PCM) subsequently steers generation toward target chemical properties. This staged design enables flexible integration of different molecular constraints within a unified generative framework. Experiments on multiple datasets demonstrate effective cross-modal controllability and strong adaptability, highlighting the potential of CMCM-DLM for heterogeneous molecular data modeling and data-driven drug discovery.
Figures & tables
Fig. 1: Overview of CMCM-DLM for adaptive cross-modality molecular generation. Starting from a pre-trained diffusion language model, Phase One introduces a structural constraint cs to establish the molecular scaffold, while Phase Two incorporates a property constraint cp to guide generation toward chemical properties. The modular controls are applied during inference without retraining the diffusion backbone.
Fig. 2: The CMCM-DLM consists of two phases illustrated in the upper part of the figure. Phase One performs single-modality control using the Structure Control Module (SCM). Phase Two enables cross-modality control by incorporating both the SCM and the Property Control Module (PCM), with their internal architectures shown in the lower part of the figure when the corresponding parameters are provided.
Fig. 3: Architecture of the scaffold-conditioned branch fθc(xt′,t,cs) in SCM. Cross-attention is introduced at each Transformer layer to incorporate the scaffold representation cs .
Basic
Property
Structure
Dataset
Method
Validity ↑
Novelty ↑
QED ↑
SAS ↓
PLogP ↑
Improvement (%) ↑
Scaffold Similarity (%) ↑
Scaffold Existence (%) ↑
GuacaMol
TGM-DLM
97%
99%
0.573
2.622
-7.013
[8%, 5%, -1%]
35%
54%
MolGPT
89%
94%
0.642
2.343
-6.872
[21%, 15%, 1%]
18%
42%
ChemFM
84%
96%
0.701
2.276
-7.366
[ 32% , 18%, -6%]
33%
39%
CMCM-DLM (Ours)
85%
97%
0.697
2.124
-6.742
[32%, 23% , 3% ]
55%
88%
ZINC
TGM-DLM
96%
98%
0.668
2.450
-4.413
[-3%, 14%, 5%]
34%
52%
TABLE I: Evaluation of controllable molecular generation methods on the GuacaMol, ZINC, and QM9 datasets. For CMCM-DLM, QED, SAS, and PLogP are jointly optimized under the scaffold constraint. “Improvement (%)” denotes the relative property improvement over the mean property values of training molecules containing the corresponding scaffold, using the absolute baseline value in the denominator. Higher QED and PLogP and lower SAS indicate better performance. The reported improvements correspond to QED, SAS, and PLogP, respectively. Scaffold membership is determined using RDKit’s ScaffoldExistence function.
Basic
Property
Structure
Dataset
Method
Validity ↑
Novelty ↑
QED ↑
SAS ↓
PLogP ↑
Improvement (%)
Scaffold Similarity (%) ↑
Scaffold Existence (%) ↑
GuacaMol
CMCM-DLM QED
80%
100%
0.710
2.610
-7.890
[34%]
51%
82%
CMCM-DLM SAS
86%
97%
0.670
2.020
-6.690
[27%]
55%
85%
CMCM-DLM PLogP
90%
99%
0.530
2.100
-5.420
[22%]
60%
87%
CMCM-DLM QED+SAS
83%
98%
0.700
2.220
-7.240
[32%, 20%]
52%
84%
CMCM-DLM QED+PLogP
82%
100%
0.690
2.300
-6.800
[30%, 2%]
52%
84%
TABLE II: Additional cross-modality control results of CMCM-DLM on GuacaMol, ZINC, and QM9 under single- and pairwise-property constraints. “Improvement (%)” reports the relative improvement of the optimized properties in the order listed in each setting. All results are evaluated on 10,000 generated molecules.
TABLE III: An optimization example from the GuacaMol dataset. The first column shows the scaffold, followed by the optimized molecules generated by CMCM-DLM under different property control combinations: QED+SAS+PLogP, QED+SAS, QED+PLogP, and SAS+PLogP. A check mark ( ✓ ) indicates improvement over the scaffold baseline, which is calculated as the mean property scores of all GuacaMol molecules containing that scaffold. The property scores (QED ( ↑ ), SAS ( ↓ ), PLogP ( ↑ )) and scaffold similarity are used to evaluate constraint satisfaction and optimization quality. The gray dashed rectangle highlights the location of the retained scaffold.
GuacaMol
Method
Validity ↑
Novelty ↑
Diversity ↑
Improvement ↑
PCM QED
86%
99%
100%
[58%]
PCM SAS
88%
96%
99%
[26%]
PCM PLogP
82%
99%
90%
[47%]
PCM HBA +
66%
100%
100%
[67%]
PCM HBA −
87%
97%
95%
[74%]
TABLE IV: PCM performance on single- and multi-property optimization using GuacaMol. The symbols + and − denote increasing and decreasing HBA/HBD, respectively. Improvement (%) measures the relative change compared with the dataset mean.
QED ↑
Method
Top-1
Top-5
Top-10
Top-100
Top-1000
GraphNVP
0.8512
0.8391
0.8260
0.7423
0.5046
GraphAF
0.9437
0.9261
0.9176
0.8593
0.5749
GDSS
0.9449
0.9369
0.9337
0.9126
0.8512
MoFlow
0.9261
0.9233
0.9150
0.8664
0.7839
D2L-OMP
0.9476
0.9417
0.9372
0.9144
0.8529
TABLE V: Comparison of QED scores for the top- k molecules out of 10,000 generated candidates. Results of GraphNVP, GraphAF, GDSS, and MoFlow are adapted from [ 15 ] .
Fig. 4: Evolution of SCM performance during training. The left plot shows results on the VS set, while the right plot shows results on five zero-shot scaffolds.
Department of Computer Science and Engineering, Shanghai Jiao Tong University, Shanghai, China · Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ), Shenzhen, Guangdong, China