Predicting spatial gene expression from histology images could scale spatial transcriptomics (ST) to image-only cohorts, but conventional histology-based ST prediction is trained and evaluated mainly by per-gene spatial-profile reconstruction. This objective is misaligned with a key downstream use of ST: differentially expressed gene (DEG) discovery, where genes are ranked for a biological or morphology-defined contrast by evidence of between-group expression differences. We formulate image-based differential expression ranking (IDER), which asks whether predicted expression profiles preserve the contrast-specific ranked gene list obtained from measured profiles. IDER compares gene rankings induced by differential-expression statistics, rather than raw expression magnitudes or per-gene spatial correlations. We further introduce a differentiable IDER objective that aligns these statistics across genes and can be trained with morphology-derived proxy contrasts without predefined biological group labels. Experiments on public ST datasets show improved DEG-ranking agreement and pathway-enrichment overlap over conventional reconstruction objectives, including morphology-derived and pathologist-annotated tissue-region evaluations.
Figures & tables
Figure 1 : (a) Conventional evaluation measures spatial-profile agreement within each gene across spots. (b) Image-based Differential Expression Ranking (IDER) ranks genes by contrast statistics for a biological or morphology-defined comparison and evaluates whether predicted profiles preserve this ranked list of DEG candidates.
Objective
Ovary
Lymph Node
Bowel
SCC DEG
nDCGDEG
SCC DEG
nDCGDEG
SCC DEG
nDCGDEG
@50
@100
@200
@50
@100
@200
@50
@100
@200
MSE [ 4 ]
0.587
0.415
0.446
0.434
0.624
0.407
0.434
0.477
0.709
0.364
0.432
0.468
PCC [ 30 ]
0.584
0.425
0.451
0.451
0.629
0.439
0.478
0.521
0.710
0.384
0.461
0.494
MSE & PCC [ 37 ]
0.609
0.437
0.465
0.453
0.650
0.477
0.511
0.548
0.714
0.383
0.459
0.496
Poisson
0.460
0.363
0.401
0.422
0.234
0.448
0.493
0.501
0.694
0.440
0.476
0.492
Table 1 : Comparison of DEG ranking performance in the cluster-based DEG setting against methods optimized with conventional objective functions.
Objective
Ovary
Lymph Node
Bowel
Breast
MSE [ 4 ]
0.326
0.382
0.336
0.144
PCC [ 30 ]
0.341
0.414
0.345
0.127
MSE & PCC [ 37 ]
0.332
0.433
0.314
0.125
Poisson
0.307
0.413
0.362
0.136
NB [ 24 ]
0.289
0.443
0.385
0.143
STRank [ 25 ]
0.302
0.383
0.360
0.147
Table 2 : Analysis of pathway enrichment overlap.
Objective
Average
adipose tissue
connective tissue
immune infiltrate
invasive cancer
SCC DEG
nDCGDEG
SCC DEG
nDCGDEG
SCC DEG
nDCGDEG
SCC DEG
nDCGDEG
SCC DEG
nDCGDEG
MSE [ 4 ]
0.239
0.084
0.393
0.069
0.206
0.080
0.197
0.176
0.272
0.082
PCC [ 30 ]
0.181
0.073
0.297
0.060
0.143
0.066
0.132
0.199
0.182
0.035
MSE & PCC [ 37 ]
0.195
0.073
0.317
0.076
0.208
0.048
0.134
0.239
0.216
0.043
Poisson
0.116
0.101
0.123
0.057
0.105
0.109
0.191
0.228
0.114
0.081
NB [ 24 ]
0.137
0.098
0.145
0.038
0.127
0.097
0.190
0.269
0.120
0.073
Table 3 : Comparison of DEG ranking performance in the real annotation-based DEG setting against methods optimized with conventional objective functions, in terms of SCC DEG and nDCGDEG @200. Due to space limitations, nDCGDEG @200 is abbreviated as nDCGDEG .
Objective
Ovary
Lymph Node
Bowel
Her2st
U stat.
DEG rank.
SCC DEG
nDCGDEG
SCC DEG
nDCGDEG
SCC DEG
nDCGDEG
SCC DEG
nDCGDEG
Baseline
✗
✗
0.584
0.451
0.629
0.521
0.710
0.494
0.181
0.073
Ours w/o DEG rank.
✓
✗
0.587
0.454
0.617
0.521
0.710
0.491
0.178
0.077
Ours
✓
✓
0.686
0.531
0.676
0.643
0.731
0.595
0.353
0.247
Table 4 : Ablation experiments of the proposed method .
Objective
Ovary
L. N.
Bowel
Breast
MSE [ 4 ]
0.235
0.171
0.345
0.113
PCC [ 30 ]
0.233
0.172
0.345
0.145
MSE & PCC [ 37 ]
0.235
0.172
0.341
0.141
Poisson
0.192
0.126
0.300
0.052
NB [ 24 ]
0.195
0.130
0.303
0.056
STRank [ 25 ]
0.190
0.130
0.300
0.065
Table 5 : Evaluation on per-gene spatial-profile PCC. L. N. indicates Lymph Node.
Objective
Ovary
L. N.
Bowel
Breast
MSE [ 4 ]
0.235
0.171
0.345
0.113
PCC [ 30 ]
0.233
0.172
0.345
0.145
MSE & PCC [ 37 ]
0.235
0.172
0.341
0.141
Poisson
0.192
0.126
0.300
0.052
NB [ 24 ]
0.195
0.130
0.303
0.056
STRank [ 25 ]
0.190
0.130
0.300
0.065
Table 5 : Evaluation on per-gene spatial-profile PCC. L. N. indicates Lymph Node.
Objective
nDCGDEG
SCC DEG
@50
@100
@200
ST-Net [ 8 ]
0.181
0.060
0.062
0.073
ST-Net + Ours
0.353
0.212
0.230
0.247
TRIPLEX [ 4 ]
0.299
0.079
0.097
0.115
TRIPLEX+ Ours
0.457
0.190
0.212
0.242
Table 6 : Plug-in experiments for existing methods on the Her2st dataset.
Spatial transcriptomics enables profiling of spatial gene expression but is limited by high cost and low throughput, motivating prediction from H&E histopathology images. Existing context-aware methods mainly supervise absolute expression, while relative expression relationships between spots are rarely used explicitly. We propose COAST, a context-aware differential learning framework for spatial gene expression prediction. COAST conditions the local and global context features with type-specific modulation and aggregates the target and context spot tokens using a Transformer encoder to capture both fine-grained local patterns and slide-level structure. It is trained with a joint objective that combines absolute expression regression with signed differential regression between the target and context spots. Experiments on multiple spatial transcriptomics datasets show consistent improvements in correlation- and distribution-based metrics, demonstrating the effectiveness of context-aware differential learning for histology-based spatial gene expression prediction.
Keunho Byeon, Sunhong Park, Jeewoo Lim +1
School of Electrical Engineering, Korea University, Seoul 02841, Republic of Korea
Predicting spatial gene expression from histopathology images enables large-scale transcriptomic profiling without the cost of direct measurement. Existing methods decode the target gene set as a flat, unstructured vector, ignoring the inter-gene dependencies arising from shared biological pathways and regulatory programs. Without explicit structural guidance, models must infer these dependencies entirely from limited paired data, constraining prediction quality. We propose MSGR (Multi-Scale Gene Refiner), which bridges this gap by incorporating the Gene Ontology (GO), a curated functional hierarchy of genes, as an explicit structural prior. MSGR organizes target genes into a four-level GO tree. Its GO-guided decoder then progressively refines predictions from coarse functional domains to fine individual genes via residual corrections under scale-weighted supervision. Operating solely on the gene side, the GO-guided decoder serves as a seamless plug-in replacement that consistently improves existing architectures without requiring any image-side modifications. Extensive experiments on nine datasets from the HEST-1k benchmark provide empirical evidence for two central claims: GO-structured decoding consistently outperforms flat decoding, even against a state-of-the-art generative baseline, and the gain is attributable to biological ontology structure rather than hierarchical decomposition per se, as confirmed by a +0.027 margin over a structurally equivalent random hierarchy.
Zhiwen Xu, Xiaoming Yan, Chengkun Wu +3
National University of Defense Technology Changsha, China
Spatial Transcriptomics (ST) has transformed biomedical research by enabling the spatial mapping of gene expression across tissue sections. However, high operational costs, specialized equipment requirements, and sensitivity to experimental noise limit the accessibility and scalability of ST. Recent computer vision approaches aim to overcome these limitations by predicting spatial gene expression directly from histopathology images. While effective, current approaches often suffer from gene expression over-smoothing and overly uniform predictions across tissue regions, suggesting that further progress depends on learning representations that reflect the hierarchical and asymmetric structure of gene regulation and tissue morphology. To address these issues, we propose Hyperbolic Contrastive Learning with Entailment for Spatial Transcriptomics (HyCLoST), a hyperbolic contrastive learning model that captures the intrinsic hierarchical relationships within ST data. By leveraging hyperbolic geometry and a gene-to-image entailment loss, HyCLoST learns structured, biologically grounded representations that improve gene expression prediction accuracy, achieving a 6% reduction in MSE and an 8% increase in PCC across 26 ST datasets, over previous methods. Our source code is publicly available at https://github.com/BCV-Uniandes/HyCLoST
Daniela Vega, Paula Cárdenas, Hannah Ceballos +2
Center for Research and Formation in Artificial Intelligence Universidad de los Andes, Colombia