Retrieval-Augmented Generation for Predicting Cellular Responses to Gene Perturbation
Authors: Andrea Giuseppe Di Francesco, Andrea Rubbi, Rishabh Jain, Pietro Liò
Organizations: Sapienza University of Rome, Rome, Italy · ISTI-CNR, Institute of Information Science and Technologies, Pisa, Italy · University of Cambridge, Cambridge, United Kingdom · Wellcome Sanger Institute, Cambridge, United Kingdom
Predicting transcriptional responses to genetic perturbations is fundamental to functional genomics and therapeutic discovery. Recent deep learning models have shown promise in single-cell perturbation response prediction, but they typically generate each response in isolation, without explicitly leveraging related perturbations. We introduce PT-RAG (Perturbation-aware Two-stage Retrieval-Augmented Generation), a plug-in retrieval-and-conditioning module for generative cellular perturbation response. PT-RAG augments an existing perturbation-response backbone with learned access to related perturbation contexts. The key challenge is that relevance is not fixed in this setting: functionally related genes may elicit different effects across cell types. PT-RAG addresses this with a two-stage retrieval mechanism: GenePT-based semantic retrieval first identifies K candidate perturbations, after which a differentiable Gumbel-Softmax selector adaptively selects retrieved contexts conditioned on the control cell state, the query perturbation, and each candidate perturbation. We evaluate PT-RAG on two backbones, a STATE-style generator used as a frozen random reservoir and a fully trained scGPT, across cross-cell-type and cross-perturbation generalization tasks. PT-RAG consistently improves distributional similarity and often overall predictive quality; for example, on scGPT cross-cell-type results, the 2-Wasserstein distance drops by 5.9% relative to scGPT alone. The code to reproduce our experiments is available at https://github.com/difra100/PT-RAG_NIPS.
Figures & tables
Figure 1: Comparison of architectures. Left : Generation baseline and Naïve RAG (non-differentiable retrieval, dotted lines; frozen components indicated by ice cube icon). Right : PT-RAG uses two-stage differentiable retrieval with Gumbel-Softmax selection conditioned on hctrl , hpert , and candidate embeddings hkcxt . The Transformer Generator is frozen in the STATE-based experiments and trained in the scGPT-based ones.
Generation
Naïve RAG
PT-RAG
Retrieval
✗
✓
✓
Cell-type-aware Retrieval
✗
✗
✓
Table 1: Method comparison. All methods except PT-RAG use non-differentiable or no retrieval; PT-RAG uses two-stage differentiable selection.
Figure 3
Metric
STATE
STATE+GenePT
Naïve RAG
PT-RAG
Gene-level expression correlations
Pearson DEG ↑
0.624†
0.631
0.396†
0.633
Spearman DEG ↑
0.403†
0.411
0.307†
0.412
Expression reconstruction accuracy
MSE PCA50 ↓
8.43
8.42
12.64†
8.39
Distributional similarity (PCA space)
Table 2: Cross-cell-type generalization. Mean values across 1,635 test perturbations from four cell types. Significance vs. PT-RAG: † ( pFDR<0.01 ), †† ( pFDR<0.05 ), ††† ( pFDR<0.1 ). Bold: best. Results with standard deviations in Table 6 .
Metric
scGPT
scGPT+GenePT
scGPT+PT-RAG
Gene-level expression correlations
Pearson DEG ↑
0.9209
0.9167
0.9128
Spearman DEG ↑
0.9500 †
0.9509 †
0.9368
Expression reconstruction accuracy
MSE PCA50 ↓
18.6440 †
19.8907 †
17.3212
Distributional similarity (PCA space)
Table 3: Cross-cell-type generalization: Results on scGPT. Mean values across 1,635 test perturbations from four cell types. Significance vs. PT-RAG: † ( pFDR<0.01 ), †† ( pFDR<0.05 ), ††† ( pFDR<0.1 ). Bold: best. Results with standard deviations in Table 7 .
Metric
STATE+GenePT
PT-RAG
Gene-level expression correlations
Pearson DEG ↑
0.6377
0.6314
Spearman DEG ↑
0.4293
0.4224
Expression reconstruction accuracy
MSE PCA50 ↓
1.1477†
1.1316
Distributional similarity (PCA space)
Table 4: Cross-perturbation generalization. Results averaged across three folds. Significance vs. PT-RAG: † ( pFDR<0.01 ), †† ( pFDR<0.05 ), ††† ( pFDR<0.1 ). Bold: best. Results with standard deviations in Table 8 .
Metric
STATE+GenePT
PT-RAG
Gene-level expression correlations
Pearson DEG ↑
0.2985 †
0.3856
Spearman DEG ↑
0.2760 †
0.3564
Expression reconstruction accuracy
MSE PCA50 ↓
5.4482
5.2871
Distributional similarity (PCA space)
Table 5: Cross-perturbation generalization on Norman et al. (CRISPRa). Mean values across the 63 test perturbations of one fold. Significance vs. PT-RAG: † ( pFDR<0.01 ), †† ( pFDR<0.05 ). Bold : best.
Figure 4: Jaccard similarity of top-10 retrieved perturbations across cell types. Off-diagonal values (range 0.185–0.196, mean 0.191) confirm that only ∼ 19% of selected perturbations overlap between cell types for the same query gene.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: PDF of Gumbel (0,1) : f(x)=e−x−e−x .
Metric
STATE
STATE+GenePT
Naïve RAG
PT-RAG
Pearson DEG ↑
0.624±0.048
0.631±0.051
0.396±0.063
0.633±0.048
Spearman DEG ↑
0.403±0.046
0.411±0.052
0.307±0.041
0.412±0.051
MSE PCA50 ↓
8.43±0.93
8.42±0.88
12.64±1.04
8.39±0.87
W1↓
35.70±1.76
35.53±1.68
48.48±1.71
35.41±1.62
W2↓
646.1±63.6
638.7±60.6
1189.5±83.3
633.7±58.4
Energy ↓
9.41±1.18
9.40±1.15
14.18±1.17
9.33±1.15
Appendix
Table 6: Cross-cell-type: full results with standard deviations. n=1,635 (375 HepG2, 416 RPE1, 443 Jurkat, 401 K562).
Metric
scGPT
scGPT+GenePT
scGPT+PT-RAG
Gene-level expression correlations
Pearson DEG ↑
0.9209 ± 0.0158
0.9167 ± 0.0149
0.9128 ± 0.0146
Spearman DEG ↑
0.9500 ± 0.0341
0.9509 ± 0.0306
0.9368 ± 0.0299
Expression reconstruction accuracy
MSE PCA50 ↓
18.6440 ± 2.0080 †
19.8907 ± 2.8476 †
17.3212 ± 2.1105
Distributional similarity (PCA space)
Appendix
Table 7: Cross-cell-type generalization on scGPT: full results with standard deviations. Mean values across 1,635 test perturbations from four cell types. Significance vs. PT-RAG: † ( pFDR<0.01 ), †† ( pFDR<0.05 ), ††† ( pFDR<0.1 ). Bold: best.
Metric
STATE+GenePT
PT-RAG
Gene-level expression correlations
Pearson DEG ↑
0.6377±0.0367
0.6314±0.0394
Spearman DEG ↑
0.4293±0.0344
0.4224±0.0381
Expression reconstruction accuracy
MSE PCA50 ↓
1.1477±0.2243†
1.1316±0.2150
Distributional similarity (PCA space)
Appendix
Table 8: Cross-perturbation generalization: full results with standard deviations. Results averaged across three folds. Significance vs. PT-RAG: † ( pFDR<0.01 ), †† ( pFDR<0.05 ), ††† ( pFDR<0.1 ). Bold: best.
Predicting cellular transcriptional responses to genetic perturbations is a central problem in single-cell biology, especially in the zero-shot setting where the perturbed gene or gene combination is unseen during training. A major difficulty is that perturbation effects are not determined by expression state alone: they depend on how the perturbed gene product influences other genes and proteins, how those downstream factors act on cis-regulatory elements, and which regulatory programs are active in the current cell state. To better capture this biological complexity, we propose CisTransCell, a cell-conditioned multi-modal framework for single-cell perturbation prediction that augments each gene with two complementary priors: a regulatory-sequence prior that captures how the gene is controlled, and a coding-sequence prior that captures what the gene product does. By integrating these priors with cellular expression state, CisTransCell models perturbation response as a cascade from gene function to regulatory control to downstream transcriptional change. Experiments on benchmark single-cell perturbation datasets show that CisTransCell achieves strong performance in zero-shot perturbation prediction.
Wei Zhang, Xun Jiang, Yuesi Xi +1
1L3S Research Center, Leibniz Universität Hannover, Germany · School of Clinical Medicine & Laboratory Medicine, Jiangsu University, China · Institute for Information Processing (tnt), Leibniz Universität Hannover, Germany.
Single-cell perturbation models can reduce costly wet-lab screening by predicting how cells respond transcriptionally to interventions. While recent generative models improve population-level prediction, individual generated cells are not explicitly checked for biological consistency. We introduce PerturbCellRL, a reinforcement learning (RL) framework that post-trains a pretrained single-cell transcriptomic generator using a suite of cell-level verifiers as rewards. These verifiers define four rewards: Pearson top-k similarity, RMSE top-k proximity, DE Spearman, and Pathway activity. The Pathway activity verifier rewards cells whose pathway responses match known perturbation biology. We evaluate PerturbCellRL on multiple genetic and chemical perturbation benchmarks. Across these benchmarks, PerturbCellRL improves over the pretrained flow-matching generator on reward-aligned evaluation metrics and a held-out evaluation metric. Moreover, PerturbCellRL remains competitive with state-of-the-art methods on population-level metrics. Together, these results frame trustworthy single-cell prediction as verifier-guided generative alignment, moving beyond matching expression distributions toward predictions whose single-cell perturbation effects are explicitly checked for biological consistency.
Dongxia Wu, Mingyu Li, Yuhui Zhang +4
Peking University Beijing, China · Stanford University Stanford, CA
Predicting transcriptomic responses to small-molecule perturbations across cell lines is central to drug discovery, but exhaustive profiling of drug-cell combinations is infeasible. We frame molecular perturbation prediction as retrieve-and-aggregate: approximate an unmeasured drug's response in a cell line by aggregating measured responses of a small set of biologically related compounds. We propose LLM-Guided Retrieval (LGR), where a large language model (LLM) ranks candidate neighbor drugs (restricted to those profiled in the target cell line); after which a fixed mean aggregator combines their observed expression deltas to form the prediction. We evaluate on the Tahoe-100M single-cell perturbation atlas under unseen-drug, unseen-cell-line, and open-world regimes. LGR consistently improves over drug mean, ChemCPA, and chemistry-based kNN baselines, with the strongest gains for unseen cell-line generalization, where it achieves higher correlation and lower error than mean baselines. Across settings, LGR improves directional (sign) accuracy of gene regulation, indicating better recovery of biologically meaningful perturbation effects even when magnitude-based metrics are similar. These results suggest that retrieval quality, rather than predictor complexity, is a key driver of zero-shot molecular perturbation prediction, and that LLMs can provide a useful biological prior when used as constrained retrieval modules.
Department of Biomedical Data Science Stanford University Stanford, CA 94305, USA · Biology Research & AI Development, Genentech DNA Way, South San Francisco, CA 94080, USA