Predicting cellular responses to perturbations is a central problem in cellular biology, with broad applications in systems biology and drug discovery. This task is challenging because cellular responses can be complex and cell-state dependent, intrinsic cell-to-cell variability can be confounded with perturbation effects, and destructive single-cell RNA sequencing precludes paired measurements of the same cell before and after treatment. Flow matching transports control cells to perturbed states flexibly, but acting on the full cell state can confound perturbation effects with pre-existing cell-to-cell variability. Disentangled approaches separate responsive from invariant components, but model perturbations through prescribed mechanisms, such as latent shifts or graph edits, limiting their flexibility. We address both limitations in a unified framework. A variational encoder disentangles each cell into an invariant block, capturing state unaffected by the perturbation, and a responsive block, capturing state it changes, through conditional priors and an information-theoretic invariance constraint. Conditional flow matching transports only the responsive block, conditioned on the perturbation and invariant state, yielding a flexible, data-driven model of perturbation effects without confounding pre-existing variability. Across several benchmarks, our method outperforms the strongest published method in settings involving combinatorial and unseen perturbation prediction.
Figures & tables
Figure 1: Overview of Drift . The two stages share one architecture. Stage 1 trains a variational autoencoder on control and perturbed cells, encoding each into an invariant code znr (driven by the covariates c : cell line, batch, sequencing depth) and a responsive latent zr (driven by the perturbation u ). Each block is associated to a learned conditional prior, p(znr∣c) and p(zr∣u) , and znr is optimized to carry no perturbation information, I(znr;u)=0 ; the concatenation [znr,zr] is decoded to reconstruct the input, control or perturbed. Stage 2 freezes the encoder and decoder and learns the mechanism p(zr∣znr,u) as a transport: a conditional flow vθ(zt,t∣znr,u) generates only the responsive latent zr′ from its control to its perturbed state along the interpolant zt ( t:0→1 ), while znr is copied from the input cell. Decoding [znr,zr′] yields the predicted perturbed cell.
Figure 2: The generative model of Drift ; shaded nodes are observed. The non-responsive block znr is generated from the assay covariates c ; the responsive block zr from the perturbation u and znr (the mechanism); the profile x from both. The non-responsive block is invariant to the perturbation, znr⊥u (Eq. 3 ), the property that defines it, while the mechanism znr→zr lets the response depend on the basal state.
Type
Method
ρΔ↑
ρΔD↑
ACCΔ↑
ACCΔD↑
DES ↑
PDS ↑
L2 ↓
MSE ↓
MAE ↓
Sta.
Control
–
–
–
–
–
0.5000
9.25
0.0487
0.1720
Linear
0.6211
0.7075
0.8007
0.8906
0.3192
0.5363
6.76
0.0265
0.1233
Linear-scGPT
0.6925
0.7514
0.8424
0.9688
0.5023
0.7913
5.31
0.0155
0.0942
Found.
scGPT
0.4408
0.4416
0.8125
0.4839
0.2504
0.5022
7.35
0.0296
0.1285
scFoundation
0.6813
0.4778
0.8768
0.7453
0.5409
0.7994
4.91
0.0138
0.0852
GeneCompass
0.6897
0.4810
0.8916
0.7484
0.5948
0.8024
4.77
0.0124
0.0808
Table 1: Norman additive split. Bold : best; underline : second best.
Type
Method
ρΔ↑
ρΔD↑
ACCΔ↑
ACCΔD↑
DES ↑
PDS ↑
L2 ↓
MSE ↓
MAE ↓
Single gene perturbation
Sta.
Control
–
–
–
–
–
0.5000
4.33
0.0108
0.0685
Linear
0.5924
0.6021
0.7695
0.8476
0.4548
0.5119
3.46
0.0067
0.0516
Linear-scGPT
0.6072
0.6352
0.7620
0.8095
0.4584
0.6571
3.60
0.0072
0.0575
Found.
scGPT
0.5172
0.5358
0.7509
0.8286
0.5987
0.5048
3.58
0.0071
0.0530
scFoundation
0.4453
0.3144
0.6882
0.7429
0.3897
0.6619
3.65
0.0079
0.0568
Table 2: Norman holdout split. Bold : best; underline : second best.
Table 4: ComboSciPlex. Bold : best; underline : second best.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Hyperparameter
Norman
Replogle
ComboSciPlex
Model
Invariant block dnr
64
Responsive block dr
192
Encoder and decoder width
1024 ( 2 layers)
Condition code dpert
128
Perturbation encoder width
256
Perturbation features
3×128
3×128
16
Appendix
Table 5: Configuration of DRIFT, read from the resolved configuration of the checkpoints that produced the tables. A single centered entry means the setting is shared by the three datasets.
probe accuracy
MINE (nats)
linear
MLP
Dataset
Conditions
chance
znr
zr
znr
zr
znr
zr
Norman additive
top 20
0.0500
0.157
0.981
0.171
0.991
0.14
3.35
top 50
0.0200
0.100
0.994
0.093
0.998
0.15
4.00
top 100
0.0100
0.053
0.980
0.056
0.981
0.17
3.96
all
0.0060
0.042
0.938
0.037
0.962
0.11
3.68
Appendix
Table 6: Decoding which perturbation a cell received from each latent block. Classes are balanced, so chance is 1/K . MINE Belghazi et al. (2018) is a mutual-information lower bound in nats.
Figure 3: Norman additive: UMAPs of the responsive block zr and invariant block znr , colored by perturbation.
Figure 4: The invariant block of Figure 3 , colored by the numeric covariates in c .
Figure 5: UMAP density visualizations on representative Norman additive perturbations.
Type
Method
Discr. Cos ↑
E-Dist ↓
Wasserstein ↓
Sta.
Control
-0.3009
13.5125
16.5615
Linear
0.3967
11.1376
15.3741
Found.
scFoundation
0.6218
1.3710
11.2452
GeneCompass
0.5623
1.3059
11.1429
Gra.
GEARS
0.5871
1.2377
11.0829
Gen.
CellFlow
0.6397
1.1979
10.7579
Appendix
Table 7: Distributional metrics on Norman additive (Appendix A ). Baselines are reported by Ruan et al. (2026) ; DRIFT is the mean and standard deviation over five seeds. Bold : best; underline : second best.
Configuration
ρΔ↑
ρΔD↑
ACCΔ↑
DES ↑
PDS ↑
rank
Norman additive
DRIFT
0.8871
0.8914
0.9222
0.9017
0.9167
1.40
w/o disentanglement
0.8588
0.8236
0.9220
0.8753
0.9042
3.40
w/o conditioning regularization
0.8974
0.8926
0.9183
0.8965
0.9051
1.80
w/o invariance penalty
0.8706
0.8778
0.9146
0.8964
0.8992
3.40
Norman holdout (single)
Appendix
Table 8: Component ablations. Each variant removes only the named component. Rank is the mean rank across five metrics (1 = best), averaged over 5 seeds.
Split
Method
ρΔ↑
ρΔD↑
ACCΔ↑
ACCΔD↑
DES ↑
PDS ↑
L2 ↓
MSE ↓
MAE ↓
Norman additive
scBIG
0.8496
0.8230
0.9197
0.9906
0.8593
0.8548
3.92
0.0091
0.0689
DRIFT , ESM2 only
0.8878
0.8938
0.9203
0.9989
0.9020
0.9045
3.22
0.0062
0.0569
DRIFT
0.8871
0.8914
0.9222
0.9977
0.9017
0.9167
3.12
0.0057
0.0549
Replogle RPE1
scBIG
0.4875
0.5925
0.6471
0.8089
0.5145
0.5520
6.12
0.0118
0.0676
DRIFT , ESM2 only
0.4919
0.6032
0.6465
0.8124
0.6865
0.5696
6.04
0.0115
0.0670
DRIFT
0.5085
0.6160
0.6537
0.8149
0.7021
0.6044
5.68
0.0101
0.0633
Appendix
Table 9: Contribution of genetic perturbation features. Drift represents each perturbed gene using ESM2 ⊕ STRING ⊕ Reactome; ESM2 only removes STRING and Reactome while leaving all other components unchanged. Ruan et al. (2026) already integrates all three sources. Bold : best; underline : second best.
Component
ρΔ
ρΔD
ACCΔ
DES
PDS
mean
Norman additive
Disentanglement
+10.6
+36.9
+0.2
+4.5
+3.3
+11.1
Conditioning regularization
− 3.9
− 0.7
+3.2
+0.9
+3.0
+0.5
Invariance penalty
+6.2
+7.4
+6.3
+0.9
+4.6
+5.1
Norman holdout (single)
Disentanglement
+40.7
+43.5
+19.0
+13.0
+2.3
+23.7
Appendix
Table 10: Component contributions, expressed as a percentage of DRIFT’s margin over the Linear baseline: (DRIFT−without component)/(DRIFT−Linear) for each dataset and metric. Green values indicate improvement; grey values lie within standard error across seeds; red values are negative.
Method
ρΔ
ρΔD
ACCΔ
ACCΔD
DES
PDS
L2
MSE
MAE
GEARS
0.0055
0.0151
0.0031
0.0080
0.0081
0.0048
0.0649
0.0006
0.0014
scGPT
0.0006
0.0045
0.0010
0.0010
0.0084
0.0026
0.0197
0.0002
0.0004
scFoundation
0.0009
0.0031
0.0004
0.0036
0.0067
0.0029
0.0044
0.0001
0.0002
CellFlow
0.0035
0.0146
0.0013
0.0213
0.0033
0.0074
0.0636
0.0004
0.0014
scBIG
0.0026
0.0038
0.0023
0.0029
0.0015
0.0096
0.0411
0.0002
0.0009
DRIFT (ours)
0.0034
0.0045
0.0031
0.0019
0.0028
0.0101
0.0558
0.0002
0.0010
Appendix
Table 11: Standard deviation over five seeds, Norman additive split.
Method
ρΔ
ρΔD
ACCΔ
ACCΔD
DES
PDS
L2
MSE
MAE
scGPT
0.0008
0.0027
0.0003
0.0031
0.0155
0.0006
0.0039
0.0001
0.0001
scFoundation
0.0072
0.0122
0.0042
0.0091
0.0158
0.0023
0.0227
0.0000
0.0005
CellFlow
0.0026
0.0098
0.0013
0.0038
0.0042
0.0092
0.0053
0.0000
0.0001
scBIG
0.0016
0.0036
0.0021
0.0014
0.0042
0.0026
0.0105
0.0000
0.0002
DRIFT (ours)
0.0048
0.0041
0.0023
0.0046
0.0061
0.0040
0.0238
0.0001
0.0004
Appendix
Table 12: Standard deviation over five seeds, Norman holdout split.
Type
Method
ρΔ↑
ρΔD↑
ACCΔ↑
ACCΔD↑
DES ↑
PDS ↑
L2 ↓
MSE ↓
MAE ↓
Gen.
DRIFT (ours)
0.5085 ± 0.0074
0.6160 ± 0.0058
0.6537 ± 0.0013
0.8149 ± 0.0015
0.7021 ± 0.0072
0.6044 ± 0.0067
5.68 ± 0.03
0.0101 ± 0.0001
0.0633 ± 0.0002
Appendix
Table 13: Replogle RPE1, mean and standard deviation over 5 seeds of DRIFT .
Method
ρΔ↑
DE-Spearman ρ↑
DS ↑
L2 ↓
MSE ↓
MAE ↓
DRIFT (ours)
0.9467 ± 0.0057
0.8901 ± 0.0057
0.8816 ± 0.0171
1.2804 ± 0.0536
0.0019 ± 0.0002
0.0150 ± 0.0006
Appendix
Table 14: ComboSciPlex, mean and standard deviation over five end-to-end seeds of Drift .
Predicting cellular transcriptional responses to genetic perturbations is a central problem in single-cell biology, especially in the zero-shot setting where the perturbed gene or gene combination is unseen during training. A major difficulty is that perturbation effects are not determined by expression state alone: they depend on how the perturbed gene product influences other genes and proteins, how those downstream factors act on cis-regulatory elements, and which regulatory programs are active in the current cell state. To better capture this biological complexity, we propose CisTransCell, a cell-conditioned multi-modal framework for single-cell perturbation prediction that augments each gene with two complementary priors: a regulatory-sequence prior that captures how the gene is controlled, and a coding-sequence prior that captures what the gene product does. By integrating these priors with cellular expression state, CisTransCell models perturbation response as a cascade from gene function to regulatory control to downstream transcriptional change. Experiments on benchmark single-cell perturbation datasets show that CisTransCell achieves strong performance in zero-shot perturbation prediction.
Wei Zhang, Xun Jiang, Yuesi Xi +1
1L3S Research Center, Leibniz Universität Hannover, Germany · School of Clinical Medicine & Laboratory Medicine, Jiangsu University, China · Institute for Information Processing (tnt), Leibniz Universität Hannover, Germany.
Virtual-cell models aim to predict how cell populations respond to perturbations, but control and treated cells are measured as unpaired populations, complicating the learning of perturbation-specific effects. We present SCALE, a conditional transport model that represents cells as unordered sets and predicts treated populations without cell-level matching. A shared set-aware encoder and conditional DiT backbone learn latent transport, making endpoint supervision directly delta-aligned without an auxiliary delta objective. Across genetic, chemical, developmental and immune perturbations, SCALE recovered gene-expression changes, response directions and population structure. In CRISPR data with dominant cell-line effects, SCALE outperformed competing methods across seven metrics and maintained separation among gene-target representations rather than collapsing them into a shared region. SCALE further prioritized cytokines predicted to produce distinct immune activation and inflammatory responses. Experiments using matched PBMC samples from three donors confirmed these predicted differences. Together, SCALE enables perturbation-specific prediction from unpaired populations and supports experimental prioritization.
Single-cell perturbation models can reduce costly wet-lab screening by predicting how cells respond transcriptionally to interventions. While recent generative models improve population-level prediction, individual generated cells are not explicitly checked for biological consistency. We introduce PerturbCellRL, a reinforcement learning (RL) framework that post-trains a pretrained single-cell transcriptomic generator using a suite of cell-level verifiers as rewards. These verifiers define four rewards: Pearson top-k similarity, RMSE top-k proximity, DE Spearman, and Pathway activity. The Pathway activity verifier rewards cells whose pathway responses match known perturbation biology. We evaluate PerturbCellRL on multiple genetic and chemical perturbation benchmarks. Across these benchmarks, PerturbCellRL improves over the pretrained flow-matching generator on reward-aligned evaluation metrics and a held-out evaluation metric. Moreover, PerturbCellRL remains competitive with state-of-the-art methods on population-level metrics. Together, these results frame trustworthy single-cell prediction as verifier-guided generative alignment, moving beyond matching expression distributions toward predictions whose single-cell perturbation effects are explicitly checked for biological consistency.
Dongxia Wu, Mingyu Li, Yuhui Zhang +4
Peking University Beijing, China · Stanford University Stanford, CA