Cross-Modal Contrastive Learning for the Retrieval of Immunotherapy-Associated Molecular Signatures from Histopathology
Authors: Sigrid Vila-Bagaria, Mar Teixidó, Miquel Piñol, Felip Vilardell, Robert Montal, Veronica Vilaplana
Organizations: Signal Theory and Communications Department (TSC) Universitat Politècnica de Catalunya - BarcelonaTech (UPC) · Cancer Biomarkers Research Group (GReBiC) - IRB Lleida
Gastric Adenocarcinoma is a leading cause of cancer mortality. Although "Inflamed/Non-Inflamed" subtypes have been proposed to predict immunotherapy response, their identification relies on a costly 10-gene RNA signature. We propose a Cross-modal Contrastive Multiple Instance Learning (CCMIL) framework for cross-modal retrieval, imputing these molecular signatures directly from standard Hematoxylin & Eosin (H&E) slides. By leveraging a supervised contrastive objective, CCMIL aligns visual morphological patterns with molecular phenotypes into a shared latent space. This establishes an interpretable search-by-case retrieval engine, enabling pathologists to query a whole slide image to surface transcriptomically coherent neighbors and approximate RNA signatures without genomic sequencing at inference. Our results demonstrate that this retrieval-first approach captures the continuous phenotypic spectrum of tumor inflammation and yields clinically interpretable attention heatmaps. Furthermore, the learned representation also supports competitive downstream classification, providing a practical molecular pre-screening strategy.
Figures & tables
Figure 1 : The CCMIL framework. (a) Dual-stream architecture aligning WSI and RNA features during training to enable inference from H&E alone. (b) The composite objective aligns modalities and separates “Inflamed/Non-Inflamed” cases in the joint latent space.
Retrieval
Retrieval (Image → RNA)
k-NN Classification
Model
Q → Tgt
R@1
R@5
MRR
Acc.
F1
Baseline
Te → Tr
–
–
.71 ± .106
.57 ± .129
.53 ± .109
Baseline
Te → Te
.67 ± .115
.89 ± .086
.78 ± .067
.67 ± .107
.65 ± .074
CCMIL
Te → Tr
–
–
.75 ± 0.049
.68 ± .080
.61 ± .138
CCMIL
Te → Te
.72 ± .085
.92 ± .051
.80 ± .067
.71 ± .155
.67 ± .133
Table 1 : Quantitative performance comparison of the proposed framework against state-of-the-art baselines. All values are reported as mean ± standard deviation across 5-fold cross-validation.
Components
Classification
Retrieval
THCA
SupConL
ClassL
Acc.
AUC
F1
R@1
R@5
✓
✓
✗
.54 ± .092
.45 ± .156
.66 ± .069
.69 ± .068
.88 ± .068
✓
✗
✓
.68 ± .103
.68 ± .073
.72 ± .076
.63 ± .119
.84 ± .068
✗
✓
✓
.72 ± .106
.74 ± .084
.74 ± .080
.60 ± .069
.84 ± .107
✓
✓
✓
.72 ± .157
.75 ± .135
.76 ± .112
.72 ± .085
.92 ± .051
Table 2 : Ablation study of the framework evaluating the Tumour Histology Conditioned Attention (THCA) and loss combinations for Classification and Retrieval.
Figure 2 : Evaluation of the proposed framework. (a) Latent space visualization via UMAP demonstrating a structured joint representation. (b) Mean RNA Cosine Similarity between the query ground truth and retrieved cases across retrieval rank k . CSLS + Reciprocal strategy remains closest to the Oracle (retrieval in the RNA space) upper bound.
Figure 3 : Qualitative evaluation of CCMIL. (a) CCMIL generates dense, highly localized attention hotspots compared to the baseline models. (b) The joint latent space captures a continuous inflammatory gradient, successfully retrieving transcriptomically coherent neighbors.
H&E-stained whole-slide images offer cohort-scale availability and rich spatial context but lack molecular specificity, whereas bulk RNA-seq provides transcriptome-wide resolution at high cost with limited archival availability. We show that training a lightweight alignment module atop frozen histopathology and RNA-Seq foundation models enables open-vocabulary molecular prompting -- querying H&E slides with gene-set signatures to predict pathway activity without sequencing or end-to-end retraining. Using contrastive learning on a multi-cancer cohort (N=1,720), we achieve a 25-fold improvement in retrieval over baseline methods. Systematic analysis reveals a graduated predictability spectrum: morphologically grounded programs (cell-cycle programs, immune-related) are most reliably predicted (R^2>0.5), while predicting pathways with no morphological footprint remains challenging as expected. We validate clinical utility on the POSEIDON clinical trial: H&E-predicted squamous cell carcinoma scores recapitulate NSCLC subtype identity and predicted IFN-gamma mirror PD-L1 tumor-cell expression groups. Furthermore, genesets describing immune activation and fibrosis predict known tumor microenvironment archetypes from histology alone. We further validate generalization of our approach across unseen cohorts and demonstrate data-efficient domain adaptation, establishing a slide-native framework for molecular analysis on H&E images.
Dominik Winter, Dominik Vonficht, Loïc Le Bescond +6
AstraZeneca Computational Pathology and Biomarkers, Munich, Germany · AstraZeneca, Early Oncology and Translation Medicine, Cambridge, UK
Identifying the Inflamed'' immunophenotype in Gastric Adenocarcinoma predicts immunotherapy response but requires an expensive 10-gene RNA signature. While deep learning on standard H\&E slides offers a scalable alternative, conventional binary classifiers oversimplify continuous RNA data and introduce label noise. To resolve this, we propose VITA (VIrtual Transcriptomic Approximation). By aligning H\&E and RNA into a joint latent space during training, VITA requires only standard H\&E at inference to retrieve morphologically similar historical cases and approximate the continuous RNA signature. Achieving 0.72 classification accuracy and a 0.66 Spearman correlation, VITA provides a cost-effective virtual transcriptomics'' pre-screening tool that preserves the continuous phenotypic spectrum without requiring genomic sequencing.
Sigrid Vila-Bagaria, Mar Teixidó, Miquel Piñol +3
Image Processing Group, Universitat Politècnica de Catalunya, Spain · Research group of Cancer Biomarkers & Oncological Pathology, IRB Lleida, Spain
Multiple instance learning (MIL) is the dominant framework for whole-slide image analysis in computational pathology, typically combining a frozen patch encoder, a projection layer, and a slide-level aggregator. While encoders and aggregators have been extensively studied, the projection layer remains a largely morphology-only bottleneck. This limits endpoints such as biomarker status and survival, which are governed by a molecular state that is not fully captured by H&E morphology. We introduce Molecularly Informed Staining Transform (MIST), a plug-in replacement for the MIL projection layer that uses paired spatial transcriptomics only during training to construct virtual molecular stains. MIST clusters gene expression profiles into cross-modal prototypes, anchors them in the frozen foundation model feature space, and uses them to reorganize H&E patch features along molecularly guided axes. It requires no transcriptomics at inference and can be inserted before standard MIL aggregators. We evaluate MIST across 23 downstream tasks and 8 MIL aggregators. MIST improves 240 of 256 configurations over the standard projection layer, with an average gain of +3.5%, observed consistently across endpoint types: +5.2% on survival prediction, +3.3% on tissue subtyping, and +2.6% on biomarker prediction. Ablations confirm that gene-derived prototypes are the primary source of the gains, while spatial, biological, and pathological analyses show that cross-modal prototype affinities capture spatially coherent molecular programs from H&E alone.
Yucheng Xing, Pei Liu, Jingying Ma +6
National University of Singapore, Singapore · Hunan University, China · Peking Union Medical College Hospital (PUMCH), China +1