Anatomy-aware Fine-grained Multimodal Fusion for Laryngopharyngeal Cancer T-Staging Prediction Using CT and Radiology Report
Authors: Xingyue Zhao, Yanzhou Su, Fang Zhang, Zhanghexuan Ji, Yirui Wang, Dazhou Guo, Sibo Ju, Yuehua Cheng, +7 more
Organizations: DAMO Academy, Alibaba Group · Department of Neurosurgery, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China · Hupan Laboratory, Hangzhou, China · Department of Radiology, Eye & ENT Hospital, Fudan University, Shanghai, China · Fuzhou University, Fuzhou, China · Department of Otolaryngology-Head & Neck Surgery, Zhongshan Hospital, Fudan University, Shanghai, China · Department of Nuclear Medicine, Chang Gung Memorial Hospital, Taiwan, ROC · Department of Radiology, Zhongshan Hospital, Fudan University, Shanghai, China · Department of Otolaryngology, Zhongshan Hospital(Xiamen), Fudan University, Xiamen, China
Accurate T-staging is crucial for guiding personalized treatment strategies for laryngopharyngeal cancer. However, current clinical practice relies on invasive biopsy procedures, whereas CT-based staging remains challenging due to the complex patterns of tumor invasion. Recent computer-aided approaches face two key challenges: 1) Structural relationship modeling: existing methods underrepresent anatomically structured patterns of tumor invasion, as they either process whole CT volumes without tumor-specific anatomical constraints or rely on labor-intensive tumor segmentation. 2) Fine-grained cross-modal alignment: while radiology reports contain organ-specific invasion details, current methods that apply global feature fusion struggle to accurately align individual anatomical structures with their corresponding textual descriptions. To address these issues, we propose an anatomy-aware multimodal framework that integrates organ-level CT context and radiology reports into a unified representation for laryngopharyngeal T-staging. The framework first constructs an Anatomy-Structured Organ Graph (AOG) that captures invasion patterns between primary sites and surrounding organs, then performs Organ-Anchored Cross-Modal Alignment (OCA) so that each organ node aggregates textual evidence from the radiology report, and finally refines this graph representation by injecting organ-specific invasion cues extracted from the report via Report-Enhanced Graph-Refinement (REG), yielding a multimodal organ graph that combines spatial and textual evidence. Extensive experiments demonstrate that the proposed framework achieves superior performance in T-staging of laryngopharyngeal cancer.
Figures & tables
Fig. 1: Comparison between (a) conventional coarse-grained image–text fusion based on global CT and report information [ 1 , 2 , 3 , 4 , 5 ] , and (b) the proposed anatomy-structured multimodal framework.
Fig. 2: Overview of the proposed framework for laryngopharyngeal cancer T-staging, integrating CT images and radiology reports through three modules: (a) Anatomy-Structured Organ Graph (AOG) for organ-level feature extraction and graph construction, (b) Organ-Anchored Cross-Modal Alignment (OCA) for fine-grained cross-modal alignment, and (c) Report-Enhanced Graph-Refinement (REG) module for structural refinement using clinical invasion information.
Symbol
Description
I={I1,I2,…,IN}
Set of N CT volumes
Ii∈RH×W×D
The i -th CT volume with D axial slices of height H and width W
R={R1,R2,…,RN}
Corresponding radiology reports
K=7
Number of predefined anatomical structures
P={1,2,3}
Indices of primary sites (supraglottic, glottic, subglottic)
T={4,5,6,7}
Indices of potentially invaded organs (esophagus, thyroid gland, trachea, oral cavity)
TABLE I: List of notations used in our framework
2-Class (Early T1–T2 vs. Late T3–T4
3-Class (T1, T2, T3–T4)
F1 (Internal)
Acc (Internal)
F1 (External)
Acc (External)
F1 (Internal)
Acc (Internal)
F1 (External)
Acc (External)
Report
ROI
Image Only
SwinViT [ 37 ]
58.82 ± 0.57
60.59 ± 1.47
43.64 ± 3.74
48.57 ± 3.54
42.02 ± 2.49
42.60 ± 2.79
33.10 ± 3.51
34.93 ± 3.15
VDPF [ 38 ]
71.47 ± 3.71
73.34 ± 3.24
61.99 ± 3.42
63.23 ± 3.42
51.53 ± 2.72
53.98 ± 1.83
42.40 ± 4.27
52.59 ± 2.15
MRSN [ 39 ]
59.68 ± 4.13
61.97 ± 4.22
49.10 ± 3.14
49.35 ± 3.16
43.57 ± 3.31
43.97 ± 3.68
34.53 ± 3.59
39.16 ± 4.54
Multimodal Methods
TABLE II: Comparison with Different Approaches for Laryngopharyngeal Cancer T-staging Classification. We evaluate different methods across binary (Early T1–T2 vs. Late T3–T4) and three-class (T1, T2, T3–T4) classification settings.
Center
T1
T2
T3
T4
Total
Split
Cases
%
FUZH
151
116
99
73
439
Training
439
66.9%
CGMH
2
4
4
15
25
Test
217
33.1%
EENT
20
29
17
21
87
TCGA [ 40 ]
24
27
18
36
105
Total
197
176
138
145
656
All
656
100%
TABLE III: Dataset statistics across four medical centers.
Two Class
Three Class
AOG
OCA
REG
F1 (macro)
Acc
F1 (macro)
Acc
58.76 ± 2.68
61.36 ± 2.27
39.16 ± 1.60
40.15 ± 1.74
✓
77.96 ± 1.23
78.41 ± 1.14
52.80 ± 2.31
53.41 ± 2.27
✓
80.97 ± 1.28
82.20 ± 1.31
58.47 ± 1.06
59.09 ± 1.14
✓
✓
81.63 ± 1.71
83.71 ± 1.74
61.52 ± 1.04
64.02 ± 1.31
✓
✓
✓
83.96 ± 0.50
84.85 ± 0.66
64.53 ± 0.70
65.53 ± 0.66
TABLE IV: Ablation Study of Key Components. Performance comparison of different component combinations for binary and three-class T-staging tasks on a representative internal fold, reported as mean ± standard deviation over three random seeds. AOG, OCA, and REG represent Anatomy-Structured Organ Graph, Organ-Anchored Cross-Modal Alignment, and Report-Enhanced Graph-Refinement, respectively.
Internal Test
External Test
Method
F1 (macro)
Acc
F1 (macro)
Acc
Global + Attention
75.45
77.27
68.47
68.66
Local + Attention
79.04
81.81
71.16
71.76
Common Attention
78.65
79.54
73.06
73.61
MultiToken Attention
69.88
71.59
68.24
68.66
Ours
81.93
82.95
76.10
76.50
TABLE V: Comparison of different attention mechanisms in our OCA on a representative internal fold. The results demonstrate the superiority of our proposed method across both datasets.
Internal Test
External Test
Method
F1 (macro)
Acc
F1 (macro)
Acc
Fully-Connected Graph
73.91
75.00
69.00
69.12
Direct Fusion
53.94
56.82
56.10
56.68
Global Context
58.26
61.36
47.02
52.07
Ours (AOG)
77.76
78.41
73.33
75.00
TABLE VI: Ablation study of anatomical topology modeling. Comparison between different approaches for modeling anatomical relationships on a representative internal fold. Fully-Connected Graph represents a complete graph connection strategy, Direct Fusion uses concatenated organ masks, and Global Context applies global feature pooling.
Fig. 3: Impact of Text Encoder Freezing. Comparison of model performance with frozen and fine-tuned text encoder on both internal validation and external testing sets. Results are shown for both binary and three-class T-staging tasks.
Fig. 4: Performance Analysis of Individual Invaded Organs. Box plots showing the distribution of (a) F1-scores and (b) Accuracy for different prediction strategies across five-fold cross-validation and external testing. Comparison includes individual invaded organs (A: Esophagus, B: Thyroid, C: Trachea, D: Oral Cavity) and global CT volume processing (E).
Fig. 5: Robustness to (a) organ segmentation noise and (b) radiology report variation, evaluated on a representative internal fold. (a) Macro-F1 under random organ-mask perturbation at ± 5% and ± 10% boundary shifts; (b) macro-F1 when each radiology report is rewritten to be less- or more-detailed while the CT images and the trained model are held fixed.
Fig. 6: Visualization of OCA cross-attention weights for three cases of different T-stages. Each bipartite graph shows organ nodes (left) connected to the most-attended report keywords (right), with line width proportional to attention weight.
Fig. 7: Representative correct predictions (top) and failure cases (bottom). Each case shows the ground truth, radiologist assessments, and our model’s prediction. The shared errors in the bottom row reflect inherent limitations of CT-based invasion assessment, detailed in Section IV-H.
Two Class
Three Class
Method
F1 (macro)
Acc
F1 (macro)
Acc
Radiologist 1
76.16
76.25
53.03
60.00
Radiologist 2
74.98
75.00
53.63
60.00
Ours
76.25
76.25
59.71
62.50
TABLE VII: Comparison of T-staging prediction performance: our method versus radiologist assessments.
Group
Invasion-token attention
Correctly staged (T1–T2)
0.058
Correctly staged (T3–T4)
0.105
Over-staged (T1–T2 → T3–T4)
0.088
Under-staged (T3–T4 → T1–T2)
0.058
TABLE VIII: Cross-modal alignment of correct and misclassified cases. Mean invasion-token attention mass placed by the potentially-invaded organ nodes on invasion-related report tokens, on the Fold 1 internal test set.
T-stage
Total
Positive invasion
Negative/absent
T1
151
20 (13.2%)
131 (86.8%)
T2
116
25 (21.6%)
91 (78.4%)
T3
99
52 (52.5%)
47 (47.5%)
T4
73
50 (68.5%)
23 (31.5%)
TABLE IX: Thyroid cartilage invasion descriptions vs. pathological T-stage across 439 FUZH reports. “Positive invasion” includes both definite and suspected cartilage invasion descriptions in CT reports.
Laryngeal cancer imaging research lacks standardised public datasets to enable reproducible deep learning (DL) model development. We present LaryngealCT, a curated benchmark of 1,029 computed tomography (CT) scans aggregated from six collections from The Cancer Imaging Archive (TCIA). Uniform 1 mm isotropic volumes of interest encompassing the larynx were extracted using a weakly supervised parameter search framework validated by clinical experts. Six 3D DL architectures (custom 3D CNN, ResNet18,50,101, DenseNet121 and MedicalNet-pretrained ResNet50) were benchmarked on (i) early (Tis,T1,T2) vs. advanced (T3,T4) and (ii) T4 vs. non-T4 classification tasks. On the independent test set, the 3D CNN achieved the strongest overall performance across global and per-class metrics (Accuracy 0.854, F1-macro 0.841) in early vs. advanced classification. In the T4 task, AU-ROC values exceeded 0.82 for most models, but sensitivity for T4 disease remained limited (less than or equal to 0.412), with ResNet101 showing the most promising calibrated T4 recall (0.706. Model explainability assessed using GradCAMpp with thyroid cartilage overlays for T4 classification task revealed anatomically plausible peri-cartilage activations, although spatial overlap was modest. Through open-source data, pretrained models, and integrated explainability tools, LaryngealCT offers a reproducible foundation for AI-driven research to support future clinical decision-making in laryngeal oncology.
Nivea Roy, Son Tran, Atul Sajjanhar +4
School of Information Technology, Deakin University, Burwood, Australia · Department of Head and Neck Surgery, Kasturba Medical College, Manipal Academy of Higher Education, Manipal, India · Department of Radiodiagnosis and Imaging, Kasturba Medical College, Manipal Academy of Higher Education, Manipal, India +1
Medical vision-language pretraining (VLP) from paired CT images and radiology reports enables scalable representation learning, but most existing methods align either whole scans with entire reports or local image regions with text fragments. These formulations underuse a key property of radiology reports: findings are organized around anatomical structures, with abnormalities described by organs, disease concepts, locations, and severity-related attributes. We propose OKA-CT, an organ-hierarchical knowledge-augmented framework for CT-report VLP. OKA-CT first converts free-text reports into organ-conditioned knowledge using radiology report parsing and LLM-assisted semantic structuring. The extracted hierarchy is used across two learning stages. Stage1 injects anatomy-grounded evidence into the CT visual representation through fine-grained organ-conditioned supervision, while Stage2 uses organ-specific report evidence to guide structured report-CT contrastive learning, where hierarchy-derived semantic soft targets treat non-paired cases with shared organ-level findings as weak semantic positives rather than uniform negatives. A lightweight query-based global branch further aggregates disease-relevant volumetric evidence for whole-scan representation. On CT-RATE and RAD-ChestCT datasets, OKA-CT achieves zero-shot abnormality diagnosis AUROCs of 84.9 and 72.2, outperforming prior CT VLP baselines. Retrieval and patch-occlusion analyses further show improved report-image alignment and stronger sensitivity to disease-associated anatomical regions.
Guoliang You, Hongming Li, Yuanwang Zhang +1
Department of Radiology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA 19104, USA · School of Engineering and Applied Science, University of Pennsylvania, Philadelphia, PA 19104, USA
Lung cancer remains one of the leading causes of cancer-related mortality worldwide. Conventional computed tomography (CT) imaging, while essential for detection and staging, has limitations in distinguishing benign from malignant lesions and providing interpretable diagnostic insights. To address this challenge, this study proposes a dual-modal artificial intelligence framework that integrates CT radiology with hematoxylin and eosin (H&E) histopathology for lung cancer diagnosis and subtype classification. The system employs convolutional neural networks to extract radiologic and histopathologic features and incorporates clinical metadata to improve robustness. Predictions from both modalities are fused using a weighted decision-level integration mechanism to classify adenocarcinoma, squamous cell carcinoma, large cell carcinoma, small cell lung cancer, and normal tissue. Explainable AI techniques including Grad-CAM, Grad-CAM++, Integrated Gradients, Occlusion, Saliency Maps, and SmoothGrad are applied to provide visual interpretability. Experimental results show strong performance with accuracy up to 0.87, AUROC above 0.97, and macro F1-score of 0.88. Grad-CAM++ achieved the highest faithfulness and localization accuracy, demonstrating strong correspondence with expert-annotated tumor regions. These results indicate that multimodal fusion of radiology and histopathology can improve diagnostic performance while maintaining model transparency, suggesting potential for future clinical decision support systems in precision oncology.
Baramee Sukumal, Aueaphum Aueawatthanaphisut
Hatyaiwittayalai School, Hat Yai, Songkhla, Thailand, 90110 · School of Information, Computer, and Communication Technology,2026 Sirindhorn International Institute of Technology, Thammasat University, Pathum Thani, Thailand