AIVC (AI Virtual Cell) is a learned simulator of cellular behavior across conditions. Predicting how a cell population responds transcriptionally to a genetic perturbation is a core task. Perturb-seq records that response by destructive sequencing, so a control cell and a perturbed cell are never observed as a pair, and cells under one condition remain heterogeneous and noisy. Regression on individual cells absorbs sampling variation into the estimated effect, whereas interpretation requires the reproducible population effect. P2P (Perturbation-to-Perturbation) takes a stochastic cell-set view as its supervision unit. Two views drawn from the same condition share a reproducible population effect and differ by view-specific variation. A permutation-invariant set encoder summarizes the control population, a structured encoder represents perturbation tokens, cellular context, dose, and combination interactions, and a gate blends empirical condition-effect memory with a neural residual. A heteroscedastic head predicts the population mean and gene-wise response variance. Under one protocol and five seeds, P2P attains the lowest expression RMSE and the highest Effect Pearson, DEG F1, and DEG average precision on each of Adamson, Norman, Replogle K562, and Replogle RPE1 relative to GenePert, LinearPert, SLIM, Scouter, and scPILOT. On Replogle K562, Effect Pearson rises from 0.643 to 0.702 and DEG F1 rises from 0.067 to 0.178 relative to Scouter, the strongest baseline on both metrics.
Figures & tables
Figure 1: Cross-view population denoising. Perturbation assays provide noisy and heterogeneous cell collections without one-to-one control pairs. P2P samples two stochastic views from each population and learns the shared condition-level effect, which is then evaluated through population expression, effect recovery, and response structure.
Figure 2: P2P architecture and readout. The permutation-invariant SetEncoder summarizes a control cell set, while the structured perturbation encoder combines perturbation tokens, context, dose, and pair interactions. A memory-gated effect head combines the empirical condition effect with a neural residual and predicts the population mean and gene-wise log variance. Target views supervise the cross-view objective during training.
Figure 3: Benchmark scale and support. The four processed files differ in cell count and condition coverage, while the RPE1 embedding shows overlapping control and perturbed states. The dashed support reference marks 80 cells per condition.
Dataset
Model
RMSE ↓
Expr. r↑
Effect r↑
DEG F1 ↑
DEG AP ↑
Direction ↑
Adamson
P2P
0.0478 ± 0.0000
0.9950 ± 0.0000
0.8894 ± 0.0000
0.5306 ± 0.0000
0.6167 ± 0.0006
0.9991 ± 0.0002
GenePert
0.0534 ± 0.0002
0.9959 ± 0.0000
0.8653 ± 0.0004
0.5031 ± 0.0000
0.5812 ± 0.0000
0.9990 ± 0.0000
LinearPert
0.0507 ± 0.0000
0.9955 ± 0.0000
0.8762 ± 0.0000
0.4950 ± 0.0000
0.5560 ± 0.0000
0.9943 ± 0.0000
SLIM
0.0496 ± 0.0000
0.9958 ± 0.0000
0.8801 ± 0.0000
0.3428 ± 0.0015
0.5546 ± 0.0000
0.9961 ± 0.0000
Scouter
0.0511 ± 0.0005
0.9954 ± 0.0001
0.8697 ± 0.0025
0.4275 ± 0.0057
0.5787 ± 0.0070
0.9973 ± 0.0000
scPILOT
0.0546 ± 0.0005
0.9947 ± 0.0001
0.8521 ± 0.0019
0.4481 ± 0.0232
0.5235 ± 0.0073
0.9956 ± 0.0003
Table 1: Main benchmark results. Values are mean plus or minus sample standard deviation across the supplied five seeds. Best values for the principal effect-recovery metrics are bold.
Figure 4: Representative gene-effect predictions. Observed and predicted effects align across Adamson EIF2S1, Norman CEBPA, Replogle K562 HSPA9, and Replogle RPE1 SMN2. The scatter views the quantity optimized by the effect-centered objective.
Variant
RMSE ↓
Expr. r
Effect r
DEG F1
DEG AP
Full P2P
0.0910
0.9829
0.6609
0.2173
0.3120
No dual-view loss
0.0987
0.9762
0.6412
0.1985
0.2211
No memory effect
0.1049
0.9694
0.6282
0.1865
0.2088
No learned gate
0.0965
0.9781
0.6473
0.2038
0.2265
No pair interaction
0.1123
0.9619
0.6096
0.1742
0.1953
No heteroscedastic head
0.0942
0.9795
0.6681
0.2097
0.2319
Table 2: RPE1 ablation. Each row removes one component from the supplied evaluation.
Figure 5: Population-space cases across the four benchmarks. Control cells, observed perturbed cells, and P2P samples are projected into a common PCA coordinate system for one representative task per dataset.
Figure 6: Graph-derived local effect structure. Observed, predicted, and residual networks are paired with module-level effects for Adamson and Norman.
Figure 7: Observed and predicted effects for representative genes. Dumbbells show the signed error for Adamson EIF2S1, Norman CEBPA, Replogle K562 HSPA9, and Replogle RPE1 SMN2.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Setting
Gaussian NLL
0.35
Effect Smooth L1
1.50
Effect correlation
0.15
Moment matching
0.08
Cross-view consistency
0.05
Crossed NLL and effect multiplier
0.25
Appendix
Table 3: Fixed objective weights and readout settings.
Figure 8: Robustness and uncertainty. The complete model is stable across the supplied seeds, component removals reduce effect and ranking quality, and predicted dispersion tracks absolute error across tasks.
Figure 9: Effect-fidelity summary. The panels report top-k overlap, direction agreement by effect magnitude, residual distributions, and the task-level Effect Pearson versus DEG AP frontier.
Figure 10: Gene-level parity and residual diagnostics for representative perturbations. The diagonal identifies exact agreement, while darker points highlight genes with large absolute effects.
Figure 11: Bland-Altman agreement analysis. The horizontal axis is the observed and predicted effect mean, and the vertical axis is the P2P minus observed effect. The solid line marks the mean difference, and dashed lines mark the 95 percent limits of agreement.
Figure 12: Cell-population reconstruction cases. The control, observed perturbation, and P2P sample clouds are shown in PCA space with their population means.
Figure 13: Combination-perturbation cases from Norman. The examples cover ETS2 plus IKZF3, CBL plus UBASH3B, CEBPB plus MAPK1, and CBL plus CNN1.
Figure 14: Ranked gene-effect comparisons for representative cases. The dot and line representation exposes the signed discrepancy for the largest observed effects.
Figure 15: Gene-level predictive intervals. Observed effects are compared with P2P means and 90 percent predictive intervals for representative genes in four datasets.
Virtual-cell models aim to predict how cell populations respond to perturbations, but control and treated cells are measured as unpaired populations, complicating the learning of perturbation-specific effects. We present SCALE, a conditional transport model that represents cells as unordered sets and predicts treated populations without cell-level matching. A shared set-aware encoder and conditional DiT backbone learn latent transport, making endpoint supervision directly delta-aligned without an auxiliary delta objective. Across genetic, chemical, developmental and immune perturbations, SCALE recovered gene-expression changes, response directions and population structure. In CRISPR data with dominant cell-line effects, SCALE outperformed competing methods across seven metrics and maintained separation among gene-target representations rather than collapsing them into a shared region. SCALE further prioritized cytokines predicted to produce distinct immune activation and inflammatory responses. Experiments using matched PBMC samples from three donors confirmed these predicted differences. Together, SCALE enables perturbation-specific prediction from unpaired populations and supports experimental prioritization.
Predicting transcriptional responses to specific perturbations is critical for understanding cellular regulatory mechanisms and accelerating drug discovery. Single-cell RNA sequencing destroys each measured cell, yielding only unpaired populations of control and perturbed cells. However, existing methods typically model perturbation prediction at the single-cell level and assume cell-to-cell correspondence, which conflicts with the unpaired nature of the observed data. To address this challenge, we propose PopPert, a framework that explicitly parameterizes population-level joint gene expression distributions for collective transcriptional state modeling. Given a control population distribution and a perturbation condition, PopPert predicts perturbation-induced changes in distribution parameters, eliminating the need for cell-level correspondence and reducing sensitivity to single-cell noise. To effectively capture gene co-expression patterns, PopPert leverages a low-rank Gaussian Copula to model cross-gene statistical dependencies and construct the joint gene expression distribution, additionally allowing sampling of synthetic perturbed single-cell profiles. Across multiple single-cell benchmarks spanning both genetic and chemical perturbations, PopPert achieves superior overall performance in differential expression recovery, perturbation effect estimation, and population-level distribution matching. These results establish population-level joint distribution learning as an effective paradigm for predicting transcriptional responses from unpaired single-cell populations. Code for PopPert is publicly available at https://github.com/whd1125/PopPert.
Handong Wang, Jiaxin Qi, Haochen Feng +1
Computer Network Information Center, Chinese Academy of Sciences · University of Chinese Academy of Sciences
Predicting single-cell transcriptional responses to genetic, chemical and cytokine perturbations is a fundamental challenge in computational biology and AI Virtual Cell (AIVC) modeling, with direct implications for drug discovery and the elucidation of gene regulatory networks. Existing approaches often rely on auxiliary cell-state encoders, hierarchical variational autoencoders, dedicated Transformer encoder-decoder modules, or gene-interaction priors to compress high-dimensional expression profiles into latent representations. While effective, these designs increase architectural complexity and may limit scalability and generalizability. This paper introduces OCOO-T, a minimalist flow-matching-based AIVC model for transcriptional perturbation response prediction. OCOO-T utilizes a vanilla Transformer stack that operates directly on continuous gene expression profiles and formulates perturbation response prediction as a continuous-time denoising process. Perturbation embeddings, dosage information, and cell-line/cell-type specificity are integrated through adaptive layer normalization and in-context tokens. Comprehensive evaluations on Tahoe100M, Replogle, and PBMC benchmarks demonstrate that OCOO-T achieves state-of-the-art performance across diverse perturbations and cell types while effectively scaling to long transcriptional profiles through patching and depatching of cellular contexts. By leveraging the simplicity of Transformer-based denoising for single-cell omics, OCOO-T provides an effective and scalable framework for in-silico cellular simulation.