RIDE: Reference-Anchored Inference-Time Diffusion Editing for Scaffold Hopping
Authors: Ruoxi Gao, Frazier N. Baker, Trieu Nguyen, Xia Ning
Organizations: Department of Computer Science and Engineering, The Ohio State University · Department of Biomedical Informatics, The Ohio State University · Translational Data Analytics Institute, The Ohio State University · Division of Medicinal Chemistry and Pharmacognosy, The Ohio State University
Scaffold hopping is a critical task in drug discovery, which seeks to discover new, structurally distinct molecules that share key functional groups and similar 3D shape with a reference binding ligand. Existing diffusion-based scaffold hopping methods formulate the problem as conditional generation of scaffolds given the functional groups. However, they lack a principled mechanism to jointly enforce 2D structural novelty and preserve the 3D shape of the reference ligand. Here, we introduce RIDE, a Reference-anchored Inference-time Diffusion Editing framework for scaffold hopping. RIDE recovers the reference diffusion noise trajectory conditioned on the binding pocket and functional groups, selects an optimal trajectory segment for editing via noise perturbation, and conducts a value-guided scaffold sampling to generate new scaffolds. Extensive experimental results demonstrate that, compared to baselines, RIDE consistently generates scaffolds with lower 2D similarity and higher 3D similarity to the reference, with an average improvements of 11.7% and 7.3%, respectively. Further analysis reveals that RIDE can accommodate various reward functions, and can preserve 3D similarity even when this is not explicitly included in the reward. Two case studies illustrate RIDE's ability to generate distinct scaffolds with different structures and properties, and its ability to introduce substantial 2D variation while maintaining very high 3D similarity. RIDE is publicly available at https://anonymous.4open.science/r/RIDE-C8A0.
Figures & tables
Figure 1: Overview of RIDE . a , RIDE generates new scaffolds based on a reference scaffold, fixed functional groups, and protein pocket; b , RIDE recovers the noise trajectory for the reference ligand from a diffusion-based molecule generation model adapted to scaffold hopping; c , RIDE identifies a reference-optimal noise trajectory segment for editing; and d , RIDE performs a value-guided transition within the reference-optimal trajectory segment for a controlled editing of the noise trajectory, thus sampling a set of new scaffolds.
Model
Similarity
Vina
QED ↑
SA ↑
Conn. (%) ↑
Sim 2D↓
Sim 3D↑
Vina S ↓
Vina M ↓
Vina D ↓
Baselines
conDitar-a
0.398
0.795
−7.192
−7.753
−8.564
0.455
0.622
85.2
IPDiff-a
0.420
0.856
−7.706
−8.176
−8.897
0.456
0.619
94.7
DiffHopp
0.426
0.812
−2.223
−6.209
−8.642
0.535
0.667
89.5
ShEPhERD
0.417
0.881
−6.471
−7.402
−8.390
0.404
0.576
29.4
RIDE(1)
conDitar-a
0.354
0.884
−7.508
−8.049
−8.833
0.427
0.601
91.5
Table 1: Comparison of RIDE with baselines on the test set with L=100 . Best and second-best results are shown in bold and underlined , respectively.
L
Similarity
Vina
QED ↑
SA ↑
Conn. (%) ↑
Sim 2D↓
Sim 3D↑
Vina S ↓
Vina M ↓
Vina D ↓
random perturbation (Eq. 11 )
T/10
RIDE(1)
0.387
0.898
−7.742
−8.213
−8.840
0.442
0.614
86.6
RIDE(2)
0.343
0.874
−7.581
−8.091
−8.813
0.425
0.604
90.0
T
RIDE(1)
0.391
0.866
−7.837
−8.308
−9.028
0.462
0.629
92.3
RIDE(2)
0.356
0.854
−7.763
−8.232
−8.915
0.450
0.620
95.2
value-guided sampling (Eq. 18 )
T/10
RIDE(1)
0.354
0.884
−7.508
−8.049
−8.833
0.427
0.601
91.5
Table 2: Comparison of random perturbation and value-guided sampling on L=T/10 and L=T (perturbation segment starting from t1∗ to the end of the diffusion trajectory), using conDitar-a as the base model. Best and second-best results are shown in bold and underlined , respectively.
Figure 2: Comparison of Sim 2D vs. Sim 3D across t1∈{nL}n=59 and t1∗ with L=T/10 . (a) conDitar-a; (b) IPDiff-a.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Inversion
Sampling
Similarity
Vina
QED ↑
SA ↑
Conn. (%) ↑
Sim 2D↓
Sim 3D↑
Vina S ↓
Vina M ↓
Vina D ↓
✓
random perturbation
RIDE(1)
0.391
0.866
−7.837
−8.308
−9.028
0.462
0.629
92.3
RIDE(2)
0.356
0.854
−7.763
−8.232
−8.915
0.450
0.620
95.2
value-guided sampling
RIDE(1)
0.340
0.882
−7.612
−8.157
−8.904
0.440
0.628
95.1
RIDE(2)
0.334
0.887
−7.539
−8.054
−8.806
0.432
0.621
93.8
✗
random perturbation
RIDE(1)
0.401
0.833
−7.668
−8.144
−8.794
0.460
0.623
91.7
Appendix
Table 3: Comparison of generation results with and without trajectory inversion using conDitar-a as the base model. Random perturbation corresponds to generation under pϕ(⋅∣t1∗) , while value-guided sampling further performs search from t1∗ to t=0 . Best and second-best results are shown in bold and underlined , respectively.
Drug discovery and development is time-consuming and resource-intensive, motivating computational approaches such as diffusion models for de novo drug design. Many such models follow the structure-based drug design (SBDD) paradigm, generating molecules to fit a target binding pocket. However, existing diffusion-based SBDD methods typically couple pocket and ligand representation learning, model interactions only at the atom level, and prioritize binding affinity over other developability properties. Here, we introduce conDitar-dev, a conditional diffusion-based SBDD framework for generating ligands with strong binding affinities and favorable ADMET properties. It consists of three modules: msPRL, a pretrained multi-scale pocket representation learning module; conDitar, a pocket-conditioned diffusion model guided by msPRL representations; and paOPT, a generation-time method for optimizing ligand developability. On a newly curated benchmark of human disease targets, conDitar outperforms state-of-the-art SBDD baselines, achieving an average binding score of -8.85 kcal/mol. Across five ADMET properties, conDitar-dev improves performance by up to 73% over conDitar. To further validate the abilities of conDitar-dev to generate developable molecules, we have applied it to two validated druggable targets: programmed death-ligand 1 (PD-L1) and colony-stimulating factor 1 receptor (CSF1R) proteins. Top-ranked generatively designed molecules and their analogs have been experimentally synthesized and biologically tested. Two molecules generated directly by conDitar-dev for PD-L1 exhibited SPR-derived KD values of 3.49 and 3.75 μM, respectively. Hit expansion based on conDitar-dev-designed molecules identified selective CSF1R inhibitors with IC50 values as low as 200 nM, while also uncovering opportunities for drug repositioning.
Ruoxi Gao, Jiangweizhi Peng, Ziqi Chen +10
Computer Science and Engineering, The Ohio State University, Columbus, OH 43210. · Industrial and System Engineering, University of Minnesota, Minneapolis, MN 55455. · Google Cloud, Google LLC, Mountain View, CA 94043. +6
We present SynLaD, a latent diffusion framework for small-molecule generation that unifies ligand-based drug design objectives (what to make) with synthetic accessibility (how to make it). Current models typically optimize one objective at the expense of the other, creating a bottleneck for discovering high-scoring and synthesizable molecules. SynLaD combines reaction-constrained generation with pharmacophore-conditioned 3D design by learning a latent space that decodes to both 3D structures and synthesis pathways. An encoder maps molecules to a latent representation used by two decoder heads: (i) a geometric head that reconstructs atom types and coordinates and (ii) an autoregressive synthesis head that outputs synthetic routes in a serialized, reaction-based notation. A diffusion transformer generates novel latents in the learned space, conditioned on pharmacophore profiles. Across analogue generation tasks for bioactive ligands, SynLaD outperforms existing baselines in synthesizable and diverse hit generation, demonstrating that a single model can produce shape-aligned molecules with feasible synthesis plans.
Miruna Cretu, John Bradshaw, Patricia Suriana +6
University of Cambridge, Cambridge, UK · Prescient Design (AI for Drug Discovery), Genentech, South San Francisco, USA · Work done during an internship at Prescient Design
Small-molecule drug discovery requires simultaneous optimization of numerous properties of candidate molecules. These properties can be investigated through the analysis of high-dimensional biological signatures, such as cell morphology and transcriptomic perturbations, which provide a rich perspective on the underlying biological mechanisms. However, existing generative methods, which use those signatures for optimization, fail to meet two key requirements: providing precise guidance toward desired phenotypic signatures while maintaining structural proximity to a known hit. We introduce PhAME (Phenotype-Aware Molecular Editing), a latent diffusion framework that overcomes this challenge by recasting molecular optimization as editing in the latent space of a pretrained graph-based VAE. Our central contribution is a compositional classifier-free guidance scheme with two independent scales, one for the phenotype-conditioning and one for similarity to the seed structure, allowing practitioners to control the tradeoff between these two objectives. Empirical evaluations across diverse benchmarks, including docking score optimization and multimodal phenotypic generation, demonstrate that PhAME achieves state-of-the-art results while maintaining high chemical validity and novelty.
Łukasz Janisiów, Sebastian Musiał, Bartosz Zieliński +2
Faculty of Mathematics and Computer Science, Jagiellonian University · Doctoral School of Exact and Natural Sciences, Jagiellonian University · Jagiellonian Center for Artificial Intelligence, Jagiellonian University +2