Predicting gene expression from H&E-stained histology images offers a scalable alternative to costly spatial transcriptomics, yet most existing methods operate at the spot level, where signals from multiple cells are aggregated and critical cellular heterogeneity is obscured. Extending this paradigm to single-cell resolution is non-trivial. Naively applying pathology foundation models faces a scale mismatch: their patch-level representations mix multiple cells, whereas per-cell cropping or resizing distorts morphology and removes local context. Conversely, segmentation-based models without strong pretrained visual encoders often lack the morphological representation capacity needed for accurate molecular prediction and inherit errors from imperfect cell boundary masks. Here, we present CELLO, an efficient end-to-end framework that performs a single pathology foundation model forward pass per image and uses grid sampling to extract location-specific features for all cells simultaneously. We further introduce a distance-decay cross-attention module that refines each cell representation using spatially biased local morphological context. Using 52 public Xenium-H&E pairs from HEST-1k that span 12 organs and approximately 10 million cells, CELLO improves the average predictive accuracy over the evaluated baselines while reducing the mean whole-slide inference time compared to DeepSpot2Cell, a 14.0x speed-up on average that excludes upstream cell segmentation. Our work establishes a scalable foundation for single-cell gene expression prediction from H&E images.
Figures & tables
Figure 1: (A) Task definition: predicting single-cell spatial gene expression from H&E images with cell locations. (B) PFM-based naive cropping methods such as DeepSpot2Cell require N per-cell forward passes and distort local morphology. (C) Cell mask based UNet3+ lacks strong pretrained visual representations. (D) CELLO performs a single PFM forward pass and uses grid sampling to extract initial feature for all cells simultaneously. A context-aware cross-attention module with 2D RoPE and Gaussian distance-decay bias captures local cellular context and cell–cell interactions.
Figure 2: Given the 2D visual token map from a single PFM forward pass, CELLO first queries a location-specific cell feature by bilinear interpolation. The cell query is then refined by distance-decay cross-attention over patch tokens, where spatially nearby tokens receive a larger prior. This produces a context-aware cell embedding that preserves precise cell-location information while incorporating local morphological context.
In-distribution (ID)
Out-of-distribution (OOD)
Tissue
Model
HVG-10
HVG-50
HVG-200
Tissue
Model
HVG-10
HVG-50
HVG-200
Lung
UNet3+
0.4293
0.2177
0.0824
Kidney
UNet3+
0.1828
0.1189
0.0367
DeepSpot2Cell
0.4082
0.2264
0.0763
DeepSpot2Cell
0.3303
0.1803
0.0610
Grid
0.5481
0.3730
0.1455
Grid
0.4728
0.3147
0.1235
CELLO
0.5537
0.3827
0.1492
CELLO
0.5515
0.2743
0.1078
Breast
UNet3+
0.3078
0.1910
0.0546
Heart
UNet3+
0.2141
0.1209
0.0430
Table 1: HVG results on the 10X-Xenium-52 benchmark. ID tissues are shown on the left and OOD tissues on the right. Higher PCC values are better. The best results within each tissue are highlighted in bold .
Figure 3: Ablation and efficiency analysis of CELLO. (A) Performance comparison of different vision encoders on ID settings. Columns show, from left to right: MPG-10, HVG-10, SVG-10. (B) Effect of CLS-token fusion. Mean Δ PCC is computed as PCC (with cls) - PCC (w/o cls), averaged across different vision encoders. (C) Total inference time over the 12 test slides, split into data loading, image-encoder forward pass, and cell features with gene head. Cell segmentation (CellViT-SAM-H), required by all methods, is shown separately. CELLO is 14.0 × faster than DeepSpot2Cell.
Figure 4: Visualization of spatial gene expression predicted from H&E. Rows show, from top to bottom: H&E image, measured gene expression, CELLO predictions, and UNet3+ predictions. Compared with UNet3+, CELLO produced substantially higher correlations with ground-truth gene expression.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Split
Organ
Health condition
# samples
Sample IDs
Train
Bowel
Cancer
3
TENX111 , TENX139 , TENX147
Train
Bowel
Healthy
1
TENX114
Train
Breast
Cancer
4
NCBI783 , TENX96 , TENX97 , TENX98
Train
Kidney
Cancer
1
TENX105
Train
Liver
Healthy
1
TENX121
Train
Lung
Cancer
1
TENX141
Appendix
Table 2: Dataset split composition of the 10X-Xenium-52. Sample IDs are grouped by data split, organ type, and health condition.
Split
Samples
Mean
Min
Median
Max
Train
36
365
280
343
480
Val
4
364
313
360
422
Test
12
414
313
400
480
Appendix
Table 3: Measured genes per sample, excluding Xenium control probes.
Split
Organ
Condition
Sample
Genes
Cells
Patches
Train
Bowel
Cancer
TENX111
425
587,016
13,816
Train
Bowel
Cancer
TENX139
480
388,015
8,858
Train
Bowel
Cancer
TENX147
422
275,875
9,997
Train
Bowel
Healthy
TENX114
325
270,788
37,902
Train
Breast
Cancer
NCBI783
280
141,829
8,177
Train
Breast
Cancer
TENX96
380
365,592
24,657
Appendix
Table 4: Per-sample statistics of the 10X-Xenium-52 (genes exclude Xenium control probes).
Figure 5: Distribution of the 10X-Xenium-52. (A) Number of samples per organ. (B) Total number of cells and image patches across organs. (C) Per-sample cell-count distribution on a log scale, with points denoting individual samples and vertical bars denoting organ-level medians. (D) Disease-state composition across organs, including cancer, diseased, and healthy samples.
In-distribution (ID)
Out-of-distribution (OOD)
Tissue
Model
M-10
M-50
M-200
S-10
S-50
S-200
Tissue
Model
M-10
M-50
M-200
S-10
S-50
S-200
Lung
UNet3+
0.4530
0.2470
0.1322
0.4530
0.2470
0.1225
Kidney
UNet3+
0.2065
0.1356
0.0697
0.2065
0.1348
0.0479
DeepSpot2Cell
0.4232
0.2580
0.1387
0.4232
0.2580
0.1279
DeepSpot2Cell
0.3617
0.1975
0.0871
0.3617
0.1970
0.0717
Grid
0.5492
0.3956
0.2390
0.5492
0.3956
0.2234
Grid
0.4879
0.3431
0.1730
0.4879
0.3431
0.1483
CELLO
0.5538
0.4060
0.2450
0.5538
0.4060
0.2286
CELLO
0.5648
0.3838
0.2085
0.5648
0.3753
0.1610
Breast
UNet3+
0.3126
0.2138
0.1048
0.3126
0.2138
0.0778
Heart
UNet3+
0.2166
0.1255
0.0600
0.2166
0.1255
0.0451
Appendix
Table 5: MPG and SVG results on the 10X-Xenium-52 benchmark. ID tissues are shown on the left and OOD tissues on the right. Higher PCC values are better. The best results within each tissue are highlighted in bold .
Figure 6: Full gene panel performance comparison between DeepSpot2Cell, UNet3+ and CELLO. (A) Full gene panel performance of ID setting. (B) Full gene panel performance of OOD setting.
Model
MPG
HVG
SVG
PCC-10 ↑
PCC-50 ↑
PCC-200 ↑
PCC-10 ↑
PCC-50 ↑
PCC-200 ↑
PCC-10 ↑
PCC-50 ↑
PCC-200 ↑
Lung
UNI
0.5451
0.3969
0.2410
0.5451
0.3753
0.1492
0.5451
0.3969
0.2253
UNI_cls
0.5441
0.3931
0.2368
0.5441
0.3714
0.1437
0.5441
0.3931
0.2204
H0
0.5533
0.4053
0.2457
0.5533
0.3795
0.1485
0.5533
0.4053
0.2292
H0_cls
0.4892
0.3734
0.2213
0.4880
0.3427
0.1249
0.4892
0.3734
0.1960
Appendix
Table 6: Ablation results on the 10X-Xenium-52 ID setting evaluated by MPG, HVG, and SVG metrics. We compare different pathology foundation models as vision encoders and investigate whether incorporating the CLS token to provide additional global contextual information improves performance. Higher values of PCC-10, PCC-50, and PCC-200 indicate better results. The best results are highlighted in bold .
ID
OOD
Setting
10
50
200
10
50
200
λ=0
0.5598
0.3882
0.1752
0.4314
0.2574
0.0975
λ=0.1
0.5565
0.3866
0.1732
0.4178
0.2473
0.0933
λ=0.25
0.5533
0.3823
0.1683
0.4147
0.2484
0.0922
λ=0.5
0.5447
0.3750
0.1665
0.4018
0.2349
0.0850
λ=1.0
0.5614
0.3808
0.1692
0.4298
0.2553
0.0950
Appendix
Table 7: Loss weight λ and distance-decay bias (50-epoch runs, HVG PCC).
Model
MPG
HVG
SVG
PCC-10 ↑
PCC-50 ↑
PCC-200 ↑
PCC-10 ↑
PCC-50 ↑
PCC-200 ↑
PCC-10 ↑
PCC-50 ↑
PCC-200 ↑
Kidney
UNI
0.5638
0.3600
0.1783
0.5220
0.3229
0.1269
0.5638
0.3600
0.1545
UNI_cls
0.5155
0.3419
0.1737
0.4850
0.3070
0.1249
0.5155
0.3419
0.1515
H0
0.5041
0.3605
0.1843
0.5041
0.3404
0.1368
0.5041
0.3604
0.1592
H0_cls
0.4927
0.3275
0.1628
0.4750
0.3045
0.1127
0.4927
0.3275
0.1378
Appendix
Table 8: Ablation results on the 10X-Xenium-52 OOD setting evaluated by MPG, HVG, and SVG metrics. We compare different pathology foundation models as vision encoders and investigate whether incorporating the CLS token to provide additional global contextual information improves performance. Higher values of PCC-10, PCC-50, and PCC-200 indicate better results. The best results are highlighted in bold .
Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few thousand genes, and applicable to only a small number of samples. H&E imaging, by contrast, is cheap and collected routinely at scale. This makes predicting single-cell expression directly from morphology a practical way to bring molecular analysis to large tissue archives. We therefore present VOICE, a multimodal foundation model that predicts single-cell gene expression from H&E images using paired Xenium data. VOICE first aligns cell centered H&E morphology from a pathology foundation model with single-cell expression embeddings from a transcriptome foundation model, trained using contrastive learning over 23 million cells. Next it predicts expression through two branches. One branch directly regresses expression from morphology. The other branch retrieves measured expression from similar reference cells, recovering genes that do not have morphological signal. Because genes vary in morphological predictability, VOICE fuses the two branches with a per-gene weight. After training, VOICE generalizes to heldout patients, slides, and partially overlapping gene panels from Xenium, and it consistently outperforms prior single-cell expression prediction methods on seven metrics.
Xin Luo, Yicheng Tao, Haoxuan Zeng +6
University of Michigan Ann Arbor, Michigan, USA · Weill Cornell Medicine New York, New York, USA
Histology-based single-cell spatial transcriptomics (ST) estimation aims to predict gene expression for individual cells from histopathological images and cell locations, reducing the need for costly single-cell ST measurements. Unlike existing histology-to-ST methods that mainly predict spot-level profiles for local regions containing multiple cells, this task requires modeling cell-to-cell expression variability, which is strongly structured by cell type. We propose Genomics-Guided Cell-Type-Specific Mixture-of-Experts (GC-MoE), which estimates cell-type probabilities with a routing network and softly combines cell-type-specific experts for gene expression prediction. To further encode cell-type-dependent gene programs, we introduce the Cell-Type-Specific Co-Expression-Aware Predictor (CAP), together with a lightweight Cell-to-Cell Interaction Attention (C2CA) module for neighboring-cell context. Experiments and ablations on public single-cell ST datasets show consistent improvements over existing single-cell and adapted spot-level baselines.
Spatial transcriptomics (ST) links gene expression with tissue morphology but remains expensive and low-throughput, motivating surrogates that infer expression from routine histology. Whole-slide H&E-to-ST inference pairs a gigapixel image with gene measurements at a sparse, irregular set of locations, making multiscale modeling challenging without incurring dense-grid overhead or quadratic token mixing. We propose HiST, a hierarchical sparse transformer that treats measured locations as a lattice-indexed sparse field and builds a dyadic encoder--decoder directly on the active tissue footprint. HiST combines sparse window attention for local geometric correspondence with resolution-changing operators for rapid multiscale context integration. For a fixed window size, the dominant runtime and memory scale with the number of observed locations rather than the dense slide area. To mitigate slide-specific acquisition variation, HiST adds a bottlenecked global conditioning pathway via a \emph{slide calibration token} that summarizes slide-level context and conditions local representations. On a multi-organ benchmark spanning diverse tissues and acquisition sources, HiST improves predictive performance over recent baselines while reducing runtime and peak memory.
Weiyi Wu, Xinwen Xu, Xingjian Diao +4
Dartmouth College · Mass General Hospital · New Jersey Institute of Technology +1