Anchor-driven Multi-modal Multi-scale Expert Selection for Survival Prediction
Authors: Tao Zhou, Ying Hu, Huazhu Fu, Yi Zhou, Xiao-Jun Wu, Haibin Ling
Organizations: School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China · Institute of Advanced Intelligence and Computing (IAIC), Agency for Science, Technology and Research (A*STAR), Singapore 138632 · School of Computer Science and Engineering, Southeast University, Nanjing 211189, China · School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi 214122, China · Westlake Intelligent Computing and Application Lab, Dept of Artificial Intelligence, Westlake University, Hangzhou, China
The integrative analysis of histopathological Whole-Slide Images (WSIs) and transcriptomic profiles holds significant promise for cancer survival prediction. However, existing methods typically project multi-modal features directly into a shared latent space without explicit alignment, leading to the entanglement of mismatched morphological cues and molecular signals. Furthermore, current fusion strategies often treat the extreme spatial heterogeneity of WSIs uniformly, lacking mechanisms to adaptively prioritize clinically relevant tissue scales for individual patients. To address these limitations, we propose an Anchor-driven Multi-modal Multi-scale Expert Selection (AM2ES) framework for survival prediction. Specifically, we present an Anchor-driven Multi-modal Fusion (AMF) module, which introduces learnable semantic anchors as cross-modal mediators to bridge the semantic gap by enforcing a structurally regularized alignment between transcriptomic features and multi-scale pathology representations. Built upon this aligned semantic space, we further design a Hierarchical Mixture-of-Experts (H-MoE) selection module to decouple the hierarchical prognostic selection process. Mimicking the pathologist's diagnostic workflow, H-MoE performs (i) Intra-scale Expert Filtering to discriminatively identify salient tumor regions within each magnification, and (ii) Inter-scale Hierarchy Routing to dynamically weight and select the most informative resolution levels. Extensive experiments on multiple TCGA cancer cohorts demonstrate that our AM2ES achieves state-of-the-art performance while offering fine-grained interpretability by visualizing how specific molecular pathways drive the expert routing decisions across tissue scales. The code will be released at https://github.com/taozh2017/AM2ES.
Figures & tables
Fig. 1 : Comparison of multimodal fusion paradigms in WSI-Omics survival analysis: (a) Early Fusion, which concatenates multi-scale pathological features prior to cross-modal interaction; (b) Hierarchical Fusion, which sequentially integrates multi-scale features layer by layer in a fixed cascading manner; and (c) Our proposed AM 2 ES framework, which explicitly decouples semantic alignment via dynamic anchors from hierarchical scale selection via H-MoE.
Fig. 2 : Overview of the proposed AM 2 ES framework for multi-modal survival prediction. The pipeline integrates multi-scale WSIs and transcriptomic profiles through two synergistic stages: (1) Dynamic Anchor-Driven Alignment: Patient-specific anchors are generated conditioned on global genomic and morphological contexts. These anchors serve as semantic mediators to align multi-modal features, where structural regularization (orthogonality and sparsity constraints) is enforced to capture diverse biological patterns. (2) Hierarchical Mixture-of-Experts (H-MoE): A two-stage mechanism handles tissue heterogeneity. The Intra-scale Expert Selection module evaluates and filters morphological signals within each resolution ( i.e. , 5×,10×,20× ) to remove redundancy. Subsequently, the Inter-scale MoE Fusion module employs rank-based alignment and feature-level gating to dynamically recalibrate and aggregate prognostic evidence across scales for final survival estimation.
Fig. 3 : Comparison of multi-modal fusion paradigms in WSI-Omics survival analysis. (a) Existing methods [ 18 ] leverage co-attention and direct concatenation to fuse cross-modal data. (b) Our anchor-driven fusion strategy acts as a semantic bridge, explicitly aligning cross-modal features via learnable anchors to construct a highly discriminative and structured representation space.
Methods
Modal
Scale
BLCA
BRCA
GBMLGG
LUAD
UCEC
Average
(N=373)
(N=955)
(N=550)
(N=452)
(N=480)
MLP [ 41 ]
g.
-
0.613 ± 0.019
0.587 ± 0.033
0.809 ± 0.029
0.617 ± 0.026
0.657 ± 0.036
0.657
SNN [ 40 ]
g.
-
0.619 ± 0.023
0.596 ± 0.027
0.805 ± 0.030
0.625 ± 0.019
0.651 ± 0.018
0.659
SNNTrans [ 42 ]
g.
-
0.627 ± 0.019
0.618 ± 0.018
0.816 ± 0.037
0.631 ± 0.023
0.641 ± 0.026
0.667
ABMIL [ 43 ]
h.
20 ×
0.622 ± 0.051
0.614 ± 0.037
0.786 ± 0.028
0.596 ± 0.067
0.649 ± 0.028
0.652
CLAM [ 39 ]
h.
20 ×
0.613 ± 0.058
0.607 ± 0.016
0.792 ± 0.032
0.590 ± 0.076
0.650 ± 0.066
0.650
TABLE I: Quantitative comparison of C-index (mean ± std) across five TCGA cohorts. Baselines are categorized by input modalities (g: genomic, h: histology) and WSI scales. The best results are in bold, and the second-best are underlined.
Fig. 4 : Kaplan-Meier survival analysis across five cancer cohorts (BLCA, BRCA, GBMLGG, LUAD, and UCEC), comparing our proposed AM 2 ES framework against Survpath and HiMT. Patients are stratified into high-risk (red) and low-risk (blue) groups based on the median predicted risk score. The shaded regions represent 95% confidence intervals.
Method
CPTAC-UCEC [ 51 ]
CPTAC-LUAD [ 52 ]
Survpath [ 49 ]
0.5034 ± 0.0118
0.5670 ± 0.0341
HiMT [ 12 ]
0.5470 ± 0.0747
0.5269 ± 0.0538
Ours
0.5651 ± 0.0690
0.5843 ± 0.0428
TABLE II: Comparison of external validation performance (C-index) on two independent CPTAC cohorts .
ID
Model Variants
BRCA
LUAD
UCEC
Avg.
(a) Impact of Input Modalities
1
g. (MLP)
0.587 ± 0.033
0.617 ± 0.026
0.657 ± 0.036
0.620
2
h. (ABMIL)
0.614 ± 0.037
0.596 ± 0.067
0.649 ± 0.028
0.620
3
h.+g. (Concat)
0.630 ± 0.021
0.632 ± 0.047
0.689 ± 0.032
0.650
(b) Impact of Multi-scale Hierarchy
4
Single Scale ( 20× )
0.663 ± 0.024
0.644 ± 0.039
0.743 ± 0.036
0.683
TABLE III : Comprehensive ablation study of AM 2 ES. The study validates: (a) input modalities, (b) multi-scale hierarchy, (c) anchor fusion mechanism, (d) hierarchical expert mechanism, and (e) different inter-scale alignment strategies .
Fig. 5 : Visualization of multimodal interpretability for glioma survival prediction. This compares a high-risk GBM patient (Left, TCGA-06-0141, survival 10.28 months) and a low-risk ODG patient (Right, TCGA-S9-A7R1, survival 169.71 months). For each patient, the visualization comprises: (1) The Whole Slide Image (WSI) overlaid with spatial attention heatmaps, highlighting regions contributing most to the prediction; (2) Representative pathological patches extracted from these high-attention regions across three distinct magnification scales (as labeled: 5× , 10× , and 20× ), capturing both macroscopic tumor architecture and microscopic cellular details; and (3) The top 25 genomic features ranked by importance using Integrated Gradients, where red and blue bars indicate positive and negative contributions to the risk prediction, respectively.
K
Dataset (C-index)
Avg.
BRCA
LUAD
UCEC
6
0.647 ± 0.023
0.660 ± 0.026
0.734 ± 0.043
0.680
8
0.669 ± 0.014
0.660 ± 0.020
0.740 ± 0.016
0.690
16
0.693 ± 0.040
0.687 ± 0.020
0.762 ± 0.028
0.714
32
0.674 ± 0.014
0.678 ± 0.029
0.748 ± 0.036
0.700
TABLE IV : Sensitivity analysis of anchor number K across LUAD, UCEC, and BRCA datasets.
K′
Dataset (C-index)
Avg.
BRCA
LUAD
UCEC
2
0.659 ± 0.036
0.653 ± 0.027
0.756 ± 0.012
0.689
4
0.693 ± 0.040
0.687 ± 0.020
0.762 ± 0.028
0.714
8
0.669 ± 0.020
0.653 ± 0.031
0.747 ± 0.024
0.690
16
0.671 ± 0.018
0.666 ± 0.031
0.744 ± 0.023
0.694
TABLE V : Sensitivity analysis of the retained top instances number K′ across BRCA, LUAD, and UCEC datasets.
Setting
Dataset
Avg.
BRCA
LUAD
UCEC
w/o Lanchor
0.678 ± 0.020
0.675 ± 0.028
0.746 ± 0.033
0.700
w/ Lanchor
0.693 ± 0.040
0.687 ± 0.020
0.762 ± 0.028
0.714
TABLE VI : Ablation on the anchor regularization loss ( Lanchor ).
Fig. 6 : Global feature importance analysis across five diverse TCGA pan-cancer datasets. The horizontal bar charts display the top-ranked genomic features (comprising both RNA-seq expression and Copy Number Variations) for each respective cohort, quantified by their overall importance to the model’s survival predictions. The identification of completely distinct, disease-specific biomarker signatures across different cohorts demonstrates the model’s robust capacity to capture highly generalized, biologically relevant molecular pathways across varied oncology domains.
Method
Params. (M)
FLOPs (G)
Peak Memory (MB)
Inference Time (ms/WSI)
MoCAT [ 8 ]
3.90
1.83
123.93
60.3
HiMT [ 12 ]
5.48
2.41
156.39
7.7
UMSA [ 13 ]
31.81
40.76
553.75
72.7
Ours
12.66
4.99
190.16
9.5
TABLE VII: Comparison of computational complexity and efficiency.
Multimodal survival prediction, a crucial yet challenging task, demands the integration of multimodal medical data (\eg Whole Slide Images (WSIs) and Genomic Profiles) to achieve accurate prognostic modeling. Given the inherent heterogeneity across modalities, the feature decoupling-fusion paradigm has emerged as a dominant approach. However, these methods have the following shortcomings: (1) fail to reduce the redundant information of modality features before decoupling, which negatively affects the feature decoupling and fusion effect;(2) lack the ability to model the fine-grained relationships of the features and capture the local information interactions between intra- and inter-modality features. To address these issues, we propose a \underline{H}ierarchical \underline{D}ecoupling-Fusion \underline{M}ixture-\underline{o}f-\underline{E}xperts (HDMoE) framework with two levels of MoE and \underline{R}andom \underline{F}eature \underline{R}eorganization (RFR) modules.In the first-level MoE, shared experts and routed experts are employed to remove redundant information and extract fine-grained specific features within each modality, while the second-level MoE facilitates fine-grained inter-modality feature decoupling. Besides, we design two RFR modules following each level of MoE to finely fuse intra- and inter-modality features, which can help the model capture more fine-grained relationships between modalities. Extensive experimental results on our private Liver Cancer (LC) and three TCGA public datasets confirm the effectiveness of our proposed method. Codes are available at https://github.com/ZJUMAI/HDMoE.
Huayi Wang, Haochao Ying, Yuyang Xu +5
Zhejiang University Hangzhou, China · Xinjiang University Urumqi, China · Hangzhou City University Hangzhou, China +1
Cancer survival prediction from whole slide images (WSIs) is a challenging task in computational pathology due to the large size, irregular shape, and high granularity of the WSIs. These characteristics make it difficult to capture the full spectrum of patterns, from subtle cellular abnormalities to complex tissue interactions, which are crucial for accurate prognosis. To address this, we propose CrossFusion, a novel multi-scale feature integration framework that extracts and fuses information from patches across different magnification levels. By effectively modeling both scale-specific patterns and their interactions, CrossFusion generates a rich feature set that enhances survival prediction accuracy. We validate our approach across six cancer types from public datasets, demonstrating significant improvements over existing state-of-the-art methods. Moreover, when coupled with domain-specific feature extraction backbones, our method shows further gains in prognostic performance compared to general-purpose backbones. The source code is available at: https://github.com/RustinS/CrossFusion
Rustin Soraki, Huayu Wang, Sitong Liu +2
University of Washington, Seattle, WA · University of California, Los Angeles, CA
We introduce ProtoPathway, an interpretable-by-design multimodal framework for cancer survival prediction that unifies whole slide imaging and transcriptomics through encoders producing biologically grounded representations on both sides of the fusion. On the histopathology side, K learnable morphological prototypes, trained end-to-end with the survival objective, serve as the slide representation itself: patches flow into prototype tokens via soft assignment, compressing variable-length patch sets into fixed task-adaptive tokens. On the genomic side, a bipartite graph neural network encodes gene expression within the Reactome pathway hierarchy, producing pathway embeddings that reflect both constituent genes and their broader biological context through bidirectional message passing over a shared gene--pathway graph. Cross-modal attention then operates over a compact prototype × pathway matrix in which prototypes query pathways, modeling the biological direction in which molecular programs give rise to tissue morphology. Because both axes carry stable task-learned identity, the attention matrix is itself an interpretability output, yielding native inference-time attribution across the full biological hierarchy, from genes through pathways and prototypes to spatial tissue maps. We evaluate on five TCGA cancer cohorts, demonstrating competitive or superior survival prediction with substantially improved biological interpretability and reduced computational cost, with interpretability claims validated through fold-stratified rank-based population-level analysis. Our source code, model weights, and Reactome pathways, together with a unified codebase reimplementing all multimodal survival baselines under identical preprocessing and evaluation, are available at: https://github.com/AmayaGS/ProtoPathway.
Amaya Gallagher-Syed, Costantino Pitzalis, Myles J. Lewis +2