Anchor-driven Multi-modal Multi-scale Expert Selection for Survival Prediction
Authors: Tao Zhou, Ying Hu, Huazhu Fu, Yi Zhou, Xiao-Jun Wu, Haibin Ling
Organizations: School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China · Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR), Singapore 138632 · School of Computer Science and Engineering, Southeast University, Nanjing 211189, China · School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi 214122, China · Westlake Intelligent Computing and Application Lab, Dept of Artificial Intelligence, Westlake University, Hangzhou, China
The integrative analysis of histopathological Whole-Slide Images (WSIs) and transcriptomic profiles holds significant promise for cancer survival prediction. However, existing methods typically project multi-modal features directly into a shared latent space without explicit alignment, leading to the entanglement of mismatched morphological cues and molecular signals. Furthermore, current fusion strategies often treat the extreme spatial heterogeneity of WSIs uniformly, lacking mechanisms to adaptively prioritize clinically relevant tissue scales for individual patients. To address these limitations, we propose an Anchor-driven Multi-modal Multi-scale Expert Selection (AM2ES) framework for survival prediction. Specifically, we present an Anchor-driven Multi-modal Fusion (AMF) module, which introduces learnable semantic anchors as cross-modal mediators to bridge the semantic gap by enforcing a structurally regularized alignment between transcriptomic features and multi-scale pathology representations. Built upon this aligned semantic space, we further design a Hierarchical Mixture-of-Experts (H-MoE) selection module to decouple the hierarchical prognostic selection process. Mimicking the pathologist's diagnostic workflow, H-MoE performs (i) Intra-scale Expert Filtering to discriminatively identify salient tumor regions within each magnification, and (ii) Inter-scale Hierarchy Routing to dynamically weight and select the most informative resolution levels. Extensive experiments on multiple TCGA cancer cohorts demonstrate that our AM2ES achieves state-of-the-art performance while offering fine-grained interpretability by visualizing how specific molecular pathways drive the expert routing decisions across tissue scales. The code will be released at https://github.com/taozh2017/AM2ES.
Figures & tables
Fig. 1 : Comparison of multimodal fusion paradigms in WSI-Omics survival analysis: (a) Early Fusion, which concatenates multi-scale pathological features prior to cross-modal interaction; (b) Hierarchical Fusion, which sequentially integrates multi-scale features layer by layer in a fixed cascading manner; and (c) Our proposed AM 2 ES framework, which explicitly decouples semantic alignment via dynamic anchors from hierarchical scale selection via H-MoE.
Fig. 2 : Overview of the proposed AM 2 ES framework for multi-modal survival prediction. The pipeline integrates multi-scale WSIs and transcriptomic profiles through two synergistic stages: (1) Dynamic Anchor-Driven Alignment: Patient-specific anchors are generated conditioned on global genomic and morphological contexts. These anchors serve as semantic mediators to align multi-modal features, where structural regularization (orthogonality and sparsity constraints) is enforced to capture diverse biological patterns. (2) Hierarchical Mixture-of-Experts (H-MoE): A two-stage mechanism handles tissue heterogeneity. The Intra-scale Expert Selection module evaluates and filters morphological signals within each resolution ( i.e. , 5×,10×,20× ) to remove redundancy. Subsequently, the Inter-scale MoE Fusion module employs rank-based alignment and feature-level gating to dynamically recalibrate and aggregate prognostic evidence across scales for final survival estimation.
Fig. 3 : Comparison of multi-modal fusion paradigms in WSI-Omics survival analysis. (a) Existing methods [ 18 ] leverage co-attention and direct concatenation to fuse cross-modal data. (b) Our anchor-driven fusion strategy acts as a semantic bridge, explicitly aligning cross-modal features via learnable anchors to construct a highly discriminative and structured representation space.
Methods
Modal
Scale
BLCA
BRCA
GBMLGG
LUAD
UCEC
Average
(N=373)
(N=955)
(N=550)
(N=452)
(N=480)
MLP [ 41 ]
g.
-
0.613 ± 0.019
0.587 ± 0.033
0.809 ± 0.029
0.617 ± 0.026
0.657 ± 0.036
0.657
SNN [ 40 ]
g.
-
0.619 ± 0.023
0.596 ± 0.027
0.805 ± 0.030
0.625 ± 0.019
0.651 ± 0.018
0.659
SNNTrans [ 42 ]
g.
-
0.627 ± 0.019
0.618 ± 0.018
0.816 ± 0.037
0.631 ± 0.023
0.641 ± 0.026
0.667
ABMIL [ 43 ]
h.
20 ×
0.622 ± 0.051
0.614 ± 0.037
0.786 ± 0.028
0.596 ± 0.067
0.649 ± 0.028
0.652
CLAM [ 39 ]
h.
20 ×
0.613 ± 0.058
0.607 ± 0.016
0.792 ± 0.032
0.590 ± 0.076
0.650 ± 0.066
0.650
TABLE I: Quantitative comparison of C-index (mean ± std) across five TCGA cohorts. Baselines are categorized by input modalities (g: genomic, h: histology) and WSI scales. The best results are in bold, and the second-best are underlined.
Fig. 4 : Kaplan-Meier survival analysis across five cancer cohorts (BLCA, BRCA, GBMLGG, LUAD, and UCEC), comparing our proposed AM 2 ES framework against Survpath and HiMT. Patients are stratified into high-risk (red) and low-risk (blue) groups based on the median predicted risk score. The shaded regions represent 95% confidence intervals.
Method
CPTAC-UCEC [ 51 ]
CPTAC-LUAD [ 52 ]
Survpath [ 49 ]
0.5034 ± 0.0118
0.5670 ± 0.0341
HiMT [ 12 ]
0.5470 ± 0.0747
0.5269 ± 0.0538
Ours
0.5651 ± 0.0690
0.5843 ± 0.0428
TABLE II: Comparison of external validation performance (C-index) on two independent CPTAC cohorts .
ID
Model Variants
BRCA
LUAD
UCEC
Avg.
(a) Impact of Input Modalities
1
g. (MLP)
0.587 ± 0.033
0.617 ± 0.026
0.657 ± 0.036
0.620
2
h. (ABMIL)
0.614 ± 0.037
0.596 ± 0.067
0.649 ± 0.028
0.620
3
h.+g. (Concat)
0.630 ± 0.021
0.632 ± 0.047
0.689 ± 0.032
0.650
(b) Impact of Multi-scale Hierarchy
4
Single Scale ( 20× )
0.663 ± 0.024
0.644 ± 0.039
0.743 ± 0.036
0.683
TABLE III : Comprehensive ablation study of AM 2 ES. The study validates: (a) input modalities, (b) multi-scale hierarchy, (c) anchor fusion mechanism, (d) hierarchical expert mechanism, and (e) different inter-scale alignment strategies .
Fig. 5 : Visualization of multimodal interpretability for glioma survival prediction. This compares a high-risk GBM patient (Left, TCGA-06-0141, survival 10.28 months) and a low-risk ODG patient (Right, TCGA-S9-A7R1, survival 169.71 months). For each patient, the visualization comprises: (1) The Whole Slide Image (WSI) overlaid with spatial attention heatmaps, highlighting regions contributing most to the prediction; (2) Representative pathological patches extracted from these high-attention regions across three distinct magnification scales (as labeled: 5× , 10× , and 20× ), capturing both macroscopic tumor architecture and microscopic cellular details; and (3) The top 25 genomic features ranked by importance using Integrated Gradients, where red and blue bars indicate positive and negative contributions to the risk prediction, respectively.
K
Dataset (C-index)
Avg.
BRCA
LUAD
UCEC
6
0.647 ± 0.023
0.660 ± 0.026
0.734 ± 0.043
0.680
8
0.669 ± 0.014
0.660 ± 0.020
0.740 ± 0.016
0.690
16
0.693 ± 0.040
0.687 ± 0.020
0.762 ± 0.028
0.714
32
0.674 ± 0.014
0.678 ± 0.029
0.748 ± 0.036
0.700
TABLE IV : Sensitivity analysis of anchor number K across LUAD, UCEC, and BRCA datasets.
K′
Dataset (C-index)
Avg.
BRCA
LUAD
UCEC
2
0.659 ± 0.036
0.653 ± 0.027
0.756 ± 0.012
0.689
4
0.693 ± 0.040
0.687 ± 0.020
0.762 ± 0.028
0.714
8
0.669 ± 0.020
0.653 ± 0.031
0.747 ± 0.024
0.690
16
0.671 ± 0.018
0.666 ± 0.031
0.744 ± 0.023
0.694
TABLE V : Sensitivity analysis of the retained top instances number K′ across BRCA, LUAD, and UCEC datasets.
Setting
Dataset
Avg.
BRCA
LUAD
UCEC
w/o Lanchor
0.678 ± 0.020
0.675 ± 0.028
0.746 ± 0.033
0.700
w/ Lanchor
0.693 ± 0.040
0.687 ± 0.020
0.762 ± 0.028
0.714
TABLE VI : Ablation on the anchor regularization loss ( Lanchor ).
Fig. 6 : Global feature importance analysis across five diverse TCGA pan-cancer datasets. The horizontal bar charts display the top-ranked genomic features (comprising both RNA-seq expression and Copy Number Variations) for each respective cohort, quantified by their overall importance to the model’s survival predictions. The identification of completely distinct, disease-specific biomarker signatures across different cohorts demonstrates the model’s robust capacity to capture highly generalized, biologically relevant molecular pathways across varied oncology domains.
Method
Params. (M)
FLOPs (G)
Peak Memory (MB)
Inference Time (ms/WSI)
MoCAT [ 8 ]
3.90
1.83
123.93
60.3
HiMT [ 12 ]
5.48
2.41
156.39
7.7
UMSA [ 13 ]
31.81
40.76
553.75
72.7
Ours
12.66
4.99
190.16
9.5
TABLE VII: Comparison of computational complexity and efficiency.