Microsatellite instability-high (MSI-H) and high tumor mutational burden (TMB-H) are clinically relevant biomarkers, yet their histopathological prediction remains challenging when models are transferred across morphologically distinct cancer types. Immune-associated spatial patterns can persist across cancers despite these morphological differences, but foundation-model-based predictors trained on a single cancer do not explicitly use this information, limiting cross-cancer generalization. To address this limitation, we propose TIRA (Tumor Immune Representation Adaptation), a target-free framework that refines frozen foundation-model representations using spatial immune topology, without requiring target-domain data during model development or test-time adaptation. TIRA uses a topology-supervised biology representation to condition tile-level attention while pooling only morphological features for joint MSI and TMB prediction. We train TIRA on TCGA-COAD+READ and evaluate it zero-shot on CPTAC-COAD, TCGA-STAD, TCGA-UCEC, and CPTAC-UCEC, covering cross-site, cross-cancer, and combined cross-cancer-site distribution shifts under UNI2, CONCH, and Virchow2. With UNI2, TIRA improved zero-shot AUROC on TCGA-STAD from 0.633 to 0.766 for MSI and from 0.651 to 0.772 for TMB. Source-derived spatial immune topology improved the cross-cancer robustness of frozen pathology foundation-model representations.
Figures & tables
Figure 1: Overview of TIRA . Stage 1 learns a topology-supervised biology representation from source data, and Stage 2 uses this representation to condition tile attention while pooling only morphological features for joint MSI/TMB prediction.
RNA immune-gene correlation
Descriptor
CD8A
CD3E
FOXP3
PRF1
PDCD1
Mean ρ
Use
Lymphocyte abundance
0.161
0.188
0.064
0.013
0.164
0.118
✓
Tumor abundance
-0.168
-0.200
-0.180
-0.174
-0.110
-0.166
×
Stromal abundance
0.171
0.146
0.192
0.146
0.124
0.156
✓
Mean lymphocyte conf.
0.173
0.201
0.073
0.027
0.175
0.130
✓
Lymphocyte variability
0.143
0.177
0.053
-0.005
0.151
0.104
✓
Table 1: Source-cohort RNA-concordance screening of 18 spatial immune descriptors in TCGA-COAD ( n=227 ). Columns report Spearman ρ with five immune genes and their mean for UNI2.
Source CV
Cross-site
Cross-cancer
Cross-cancer + site
FM
Method
COAD+READ ( n =443, MSI-H=64)
CPTAC-COAD ( n =75, MSI-H=15)
TCGA-STAD ( n =308, MSI-H=54)
TCGA-UCEC ( n =297, MSI-H=105)
CPTAC-UCEC ( n =95, MSI-H=25)
UNI2
ABMIL
0.8778 (0.8244–0.9244)
0.8022 (0.6628–0.9184)
0.6333 (0.5458–0.7210)
0.5153 (0.4437–0.5870)
0.4314 (0.2868–0.5822)
CLAM-SB
0.8715 (0.8136–0.9232)
0.7089 (0.5675–0.8364)
0.6490 (0.5628–0.7326)
0.5061 (0.4358–0.5771)
0.4829 (0.3414–0.6260)
TransMIL
0.8861 (0.8369–0.9289)
0.7889 (0.6487–0.9100)
0.6395 (0.5489–0.7281)
0.5215 (0.4486–0.5937)
0.5069 (0.3721–0.6401)
CasNet-FM
0.8543 (0.7871–0.9151)
0.7489 (0.6094–0.8712)
0.6492 (0.5632–0.7340)
0.5116 (0.4427–0.5823)
0.4543 (0.3161–0.5948)
ILRA
0.8361 (0.7675–0.8975)
0.4489 (0.2804–0.6195)
0.6457 (0.5550–0.7324)
0.4887 (0.4187–0.5605)
0.5023 (0.3673–0.6370)
Table 2: Patient-level MSI prediction AUROC (95% bootstrap CI) for models trained on TCGA-COAD+READ and evaluated zero-shot on external cohorts spanning cross-site, cross-cancer, and combined shift.
Source CV
Cross-site
Cross-cancer
Cross-cancer + site
FM
Method
COAD+READ ( n =428, TMB-H=62)
CPTAC-COAD ( n =75, TMB-H=15)
TCGA-STAD ( n =308, TMB-H=56)
TCGA-UCEC ( n =213, TMB-H=36)
CPTAC-UCEC ( n =95, TMB-H=32)
UNI2
ABMIL
0.9180 (0.8721–0.9552)
0.8244 (0.6862–0.9385)
0.6509 (0.5658–0.7352)
0.5279 (0.4275–0.6275)
0.4807 (0.3485–0.6097)
CLAM-SB
0.9040 (0.8604–0.9414)
0.7900 (0.6413–0.9176)
0.6503 (0.5623–0.7365)
0.5367 (0.4389–0.6347)
0.5069 (0.3750–0.6371)
TransMIL
0.8832 (0.8329–0.9279)
0.8300 (0.7078–0.9289)
0.6451 (0.5558–0.7334)
0.5323 (0.4282–0.6347)
0.5184 (0.3883–0.6462)
CasNet-FM
0.9149 (0.8636–0.9577)
0.8200 (0.6902–0.9272)
0.6698 (0.5865–0.7509)
0.5430 (0.4436–0.6419)
0.4742 (0.3421–0.6057)
ILRA
0.8623 (0.8032–0.9138)
0.5189 (0.3561–0.6799)
0.6787 (0.5955–0.7580)
0.4735 (0.3725–0.5741)
0.4747 (0.3486–0.6035)
Table 3: Patient-level TMB prediction AUROC (95% bootstrap CI) for models trained on TCGA-COAD+READ and evaluated zero-shot on external cohorts spanning cross-site, cross-cancer, and combined shift.
MSI
TMB
Source
Target
Shift
ABMIL
TIRA
Δ AUROC [95% CI]
ABMIL
TIRA
Δ AUROC [95% CI]
COAD+READ
TCGA-STAD
Cross-cancer
0.633
0.766
+0.133 [0.076, 0.193]
0.651
0.772
+0.121 [0.063, 0.182]
COAD+READ
TCGA-UCEC
Cross-cancer
0.515
0.595
+0.080 [0.029, 0.131]
0.528
0.587
+0.059 [ − 0.026, 0.143]
COAD+READ
CPTAC-UCEC
Cross-cancer + site
0.431
0.561
+0.130 [ − 0.032, 0.297]
0.481
0.532
+0.051 [ − 0.081, 0.185]
STAD+READ
TCGA-COAD
Cross-cancer
0.786
0.859
+0.073 [0.015, 0.137]
0.805
0.889
+0.084 [0.028, 0.145]
STAD+READ
TCGA-UCEC
Cross-cancer
0.519
0.678
+0.159 [0.094, 0.226]
0.550
0.646
+0.096 [ − 0.014, 0.204]
Table 4: Reciprocal robustness analysis under UNI2: patient-level AUROC for matched ABMIL and TIRA with paired 95% CIs. Descriptor identities are fixed from the primary TCGA-COAD RNA-concordance screening across all configurations.
Source CV
Cross-site
Cross-cancer
Cross-cancer + site
Gen. Drop ↓
CC Drop from Full
Model
COAD+READ
CPTAC-COAD
TCGA-STAD
TCGA-UCEC
CPTAC-UCEC
Source → CC
Mean STAD/UCEC
MSI
TMB
MSI
TMB
MSI
TMB
MSI
TMB
MSI
TMB
MSI
TMB
MSI
TMB
ABMIL
0.878
0.918
0.802
0.824
0.633
0.651
0.515
0.528
0.431
0.481
0.304
0.329
0.106
0.090
TIRA Full
0.893
0.922
0.793
0.843
0.766
0.772
0.595
0.587
0.561
0.532
0.212
0.242
–
–
− Biology-Guided Attention
0.876
0.913
0.782
0.802
0.665
0.673
0.526
0.525
0.533
0.554
0.281
0.314
0.085
0.081
− Direction Suppression
0.891
0.918
0.823
0.841
0.682
0.690
0.520
0.540
0.498
0.523
0.290
0.303
0.080
0.065
Table 5: Component ablation of TIRA (UNI2, COAD+READ source): patient-level AUROC across source and zero-shot cohorts, with CC Drop from Full denoting the reduction in mean STAD/UCEC AUROC relative to full TIRA.
Dataset
Task
No Supp.
Random
READ
Δ READ–Random
TCGA-STAD
MSI
0.682
0.712
0.766
+0.054 [+0.010,+0.100]
TCGA-STAD
TMB
0.690
0.724
0.772
+0.048 [+0.000,+0.090]
TCGA-UCEC
MSI
0.520
0.553
0.595
+0.042 [+0.001,+0.055]
TCGA-UCEC
TMB
0.540
0.521
0.587
+0.066 [+0.001,+0.145]
Table 6: Reference-choice ablation under UNI2. READ is compared with matched random-COAD and no-suppression conditions; Δ and 95% CIs report READ versus random-COAD using 10,000 paired bootstrap resamples.
Figure 2: Representative UNI2 attention maps for MSI-H and TMB-H cases from TCGA-STAD and TCGA-UCEC. Tiles are colored by log10 attention weight; cyan outlines denote tumor–lymphocyte interface tiles.
Table 8: Raw probability calibration for TIRA and ABMIL on UNI2 zero-shot target cohorts.
Descriptor
CD8A
CD3E
FOXP3
PRF1
PDCD1
Mean ρ
Pass
Lymphocyte abundance
0.184
0.234
0.136
0.113
0.249
0.183
✓
Tumor abundance
-0.049
-0.065
-0.098
-0.024
0.004
-0.047
×
Stromal abundance
0.063
0.053
0.119
0.049
0.028
0.062
×
Mean lymphocyte conf.
0.188
0.240
0.148
0.119
0.257
0.190
✓
Lymphocyte variability
0.183
0.232
0.122
0.098
0.241
0.175
✓
Local immune nbhd.
0.098
0.129
0.106
0.024
0.082
0.088
×
Table S1: RNA-concordance for CONCH in TCGA-COAD ( n=227 ).
Descriptor
CD8A
CD3E
FOXP3
PRF1
PDCD1
Mean ρ
Pass
Lymphocyte abundance
0.181
0.185
0.062
0.065
0.186
0.136
✓
Tumor abundance
-0.101
-0.132
-0.168
-0.055
-0.038
-0.099
×
Stromal abundance
0.015
-0.010
0.050
-0.008
0.027
0.015
×
Mean lymphocyte conf.
0.205
0.202
0.074
0.076
0.205
0.153
✓
Lymphocyte variability
0.160
0.170
0.040
0.036
0.168
0.115
✓
Local immune nbhd.
0.181
0.208
0.157
0.068
0.123
0.147
✓
Table S2: RNA-concordance for Virchow2 in TCGA-COAD ( n=227 ).
Figure S1: Zero-shot calibration of ABMIL and TIRA under UNI2 on TCGA-STAD and TCGA-UCEC. Curves use five quantile-based bins; the dashed diagonal denotes perfect calibration.
Figure S2: Decision-curve analysis of ABMIL and TIRA under UNI2 on TCGA-STAD and TCGA-UCEC for MSI and TMB. Net benefit is shown across threshold probabilities from 0.01 to 0.50, with treat-all and treat-none strategies included as references.
FM
Model
COAD+READ
CPTAC-COAD
TCGA-STAD
TCGA-UCEC
CPTAC-UCEC
UNI2
ABMIL
0.823
0.500
0.566
0.496
0.477
CLAM-SB
0.833
0.492
0.582
0.513
0.511
TransMIL
0.814
0.608
0.566
0.514
0.543
CasNet-FM
0.844
0.617
0.544
0.510
0.489
ILRA
0.798
0.467
0.598
0.512
0.501
TIRA (Ours)
0.841
0.742
0.647
0.531
0.540
Table S3: Balanced accuracy for MSI prediction across foundation models and evaluation cohorts.
FM
Model
COAD+READ
CPTAC-COAD
TCGA-STAD
TCGA-UCEC
CPTAC-UCEC
UNI2
ABMIL
0.866
0.625
0.557
0.473
0.501
CLAM-SB
0.833
0.742
0.571
0.531
0.485
TransMIL
0.815
0.708
0.565
0.543
0.494
CasNet-FM
0.881
0.558
0.574
0.468
0.478
ILRA
0.813
0.550
0.635
0.457
0.492
TIRA (Ours)
0.876
0.734
0.640
0.538
0.499
Table S4: Balanced accuracy for TMB prediction across foundation models and evaluation cohorts.
Predicting microsatellite instability (MSI) status from routine hematoxylin and eosin (H&E) whole slide images (WSIs) offers a practical alternative to molecular testing, but models trained at one institution tend to generalize poorly to slides acquired at a different site. Foundation model representations, despite their generality, still encode site-specific texture alongside the conserved biological morphology underlying MSI. We investigate whether tile-level spatial priors derived from known MSI histology can guide these representations toward more site-invariant features. We introduce a biologically motivated spatial prior based on peripheral distance encoding, reflecting the Crohn's-like peripheral lymphocytic reaction at the tumor invasive margin, and evaluate a secondary local immune neighborhood encoding reflecting the lymphocyte-to-tumor ratio in each tile's immediate spatial neighborhood. Both priors are injected into a TransMIL aggregator before self-attention, allowing the transformer to integrate spatial biological context with UNI2-h or Virchow2 features across all attention layers. We evaluate six foundation model and MIL aggregator combinations as a reference, then assess the effect of each spatial prior. Training on TCGA-COAD (137 slides) and evaluating externally on TCGA-READ (50 slides) without retraining, peripheral distance encoding achieves MSI AUC 0.959 +/- 0.012 on COAD and MSS specificity 1.000 on READ, compared to 0.957 and 0.939 for the strongest reference configuration. Local immune neighborhood encoding achieves comparable internal AUC but lower cross-site specificity, suggesting margin proximity encodes a more site-invariant biological signal than local immune density. Results suggest biologically grounded spatial priors act as regularizers that reduce reliance on site-specific imaging patterns.
Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under distribution shift remains underexplored. This work systematically evaluates FM-based representations on a suite of computational pathology tasks across two real-world commercial cohorts, IH-BC and IH-NSCLC, drawn from the licensed in-house (IH) oncology dataset. The analysis focuses on two modalities, whole-slide images and transcriptomic profiles, drawn from the IH multimodal data. We first benchmark unimodal probing performance across five FMs on eight downstream classification tasks, and find that image and omics representations carry complementary predictive signals. Then we investigate whether multimodal fusion can yield additional gains over unimodal baselines by comparing three image-omics fusion strategies built on paired representations. The trustworthiness of selected unimodal and multimodal pipelines is further assessed through conformal prediction. Our results show that FM representations achieve competitive performance on out-of-distribution data and that multimodal fusion helps mainly when no single modality dominates the signal. Conformal prediction reveals that in the majority of cases where a point prediction fails, the true diagnosis remains recoverable within the prediction set, reinforcing the value of uncertainty-aware inference for clinical support.
Jingyu Hu, Giuseppe Tripodi, Reed Naidoo +2
The Alan Turing Institute, London, United Kingdom · University of Manchester, Manchester, United Kingdom · The Institute of Cancer Research, London, United Kingdom +1
Predicting immune biomarkers associated with the tumor immune microenvironment (TIME) is critical for advancing precision oncology, yet existing approaches are largely limited to single image modalities and suffer from insufficient resolution and incomplete utilization of complementary clinical and biological information. Here we introduce MixTIME, a multimodal foundation model that leverages a mixture-of-experts (MoE) architecture to integrate pathology foundation models trained across distinct modalities: image only (UNIv2), image text (CONCHv1.5), and image transcriptomic (STPath) representations for pixel-level and slide-level prediction of multiplex immunofluorescence (mIF) protein expression from hematoxylin and eosin (HE) whole-slide images. MixTIME employs a learnable router to dynamically weight expert contributions and is trained with a distribution- and tendency-aware loss function. Benchmarked on two datasets of different scales, MixTIME achieves state-of-the-art performance across 17 protein markers as measured by correlation metrics. The predicted mIF profiles substantially enhance downstream tasks, including spatial domain identification, survival prediction, and AI-assisted pathology report generation validated by expert pathologists from multiple institutes across the world. Furthermore, MixTIME enables longitudinal tracking of protein expression dynamics across clinical time points and reveals protein gene interaction patterns linked to drug resistance and immune suppression in tumor microenvironments. Collectively, MixTIME provides a scalable framework for multimodal biomarker discovery and clinical translation in computational pathology.
Tianyu Liu, Ziqing Wang, Zhaokang Liang +12
Program of Computational Biology and Bioinforamtics, Yale University, USA · Broad Institute of MIT and Harvard, USA · These authors contributed equally to this work. +12