Anatomy-aware Fine-grained Multimodal Fusion for Laryngopharyngeal Cancer T-Staging Prediction Using CT and Radiology Report
Organizations: DAMO Academy, Alibaba Group · Department of Neurosurgery, Peking Union Medical College Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China · Hupan Laboratory, Hangzhou, China · Department of Radiology, Eye & ENT Hospital, Fudan University, Shanghai, China · Fuzhou University, Fuzhou, China · Department of Otolaryngology-Head & Neck Surgery, Zhongshan Hospital, Fudan University, Shanghai, China · Department of Nuclear Medicine, Chang Gung Memorial Hospital, Taiwan, ROC · Department of Radiology, Zhongshan Hospital, Fudan University, Shanghai, China · Department of Otolaryngology, Zhongshan Hospital(Xiamen), Fudan University, Xiamen, China
Abstract
Accurate T-staging is crucial for guiding personalized treatment strategies for laryngopharyngeal cancer. However, current clinical practice relies on invasive biopsy procedures, whereas CT-based staging remains challenging due to the complex patterns of tumor invasion. Recent computer-aided approaches face two key challenges: 1) Structural relationship modeling: existing methods underrepresent anatomically structured patterns of tumor invasion, as they either process whole CT volumes without tumor-specific anatomical constraints or rely on labor-intensive tumor segmentation. 2) Fine-grained cross-modal alignment: while radiology reports contain organ-specific invasion details, current methods that apply global feature fusion struggle to accurately align individual anatomical structures with their corresponding textual descriptions. To address these issues, we propose an anatomy-aware multimodal framework that integrates organ-level CT context and radiology reports into a unified representation for laryngopharyngeal T-staging. The framework first constructs an Anatomy-Structured Organ Graph (AOG) that captures invasion patterns between primary sites and surrounding organs, then performs Organ-Anchored Cross-Modal Alignment (OCA) so that each organ node aggregates textual evidence from the radiology report, and finally refines this graph representation by injecting organ-specific invasion cues extracted from the report via Report-Enhanced Graph-Refinement (REG), yielding a multimodal organ graph that combines spatial and textual evidence. Extensive experiments demonstrate that the proposed framework achieves superior performance in T-staging of laryngopharyngeal cancer.
Figures & tables
| Symbol | Description |
| Set of CT volumes | |
| The -th CT volume with axial slices of height and width | |
| Corresponding radiology reports | |
| Number of predefined anatomical structures | |
| Indices of primary sites (supraglottic, glottic, subglottic) | |
| Indices of potentially invaded organs (esophagus, thyroid gland, trachea, oral cavity) |
| 2-Class (Early T1–T2 vs. Late T3–T4 | 3-Class (T1, T2, T3–T4) | |||||||||
| F1 (Internal) | Acc (Internal) | F1 (External) | Acc (External) | F1 (Internal) | Acc (Internal) | F1 (External) | Acc (External) | Report | ROI | |
| Image Only | ||||||||||
| SwinViT [ 37 ] | 58.82 0.57 | 60.59 1.47 | 43.64 3.74 | 48.57 3.54 | 42.02 2.49 | 42.60 2.79 | 33.10 3.51 | 34.93 3.15 | ||
| VDPF [ 38 ] | 71.47 3.71 | 73.34 3.24 | 61.99 3.42 | 63.23 3.42 | 51.53 2.72 | 53.98 1.83 | 42.40 4.27 | 52.59 2.15 | ||
| MRSN [ 39 ] | 59.68 4.13 | 61.97 4.22 | 49.10 3.14 | 49.35 3.16 | 43.57 3.31 | 43.97 3.68 | 34.53 3.59 | 39.16 4.54 | ||
| Multimodal Methods | ||||||||||
| Center | T1 | T2 | T3 | T4 | Total | Split | Cases | % |
| FUZH | 151 | 116 | 99 | 73 | 439 | Training | 439 | 66.9% |
| CGMH | 2 | 4 | 4 | 15 | 25 | Test | 217 | 33.1% |
| EENT | 20 | 29 | 17 | 21 | 87 | |||
| TCGA [ 40 ] | 24 | 27 | 18 | 36 | 105 | |||
| Total | 197 | 176 | 138 | 145 | 656 | All | 656 | 100% |
| Two Class | Three Class | |||||
| AOG | OCA | REG | F1 (macro) | Acc | F1 (macro) | Acc |
| 58.76 2.68 | 61.36 2.27 | 39.16 1.60 | 40.15 1.74 | |||
| ✓ | 77.96 1.23 | 78.41 1.14 | 52.80 2.31 | 53.41 2.27 | ||
| ✓ | 80.97 1.28 | 82.20 1.31 | 58.47 1.06 | 59.09 1.14 | ||
| ✓ | ✓ | 81.63 1.71 | 83.71 1.74 | 61.52 1.04 | 64.02 1.31 | |
| ✓ | ✓ | ✓ | 83.96 0.50 | 84.85 0.66 | 64.53 0.70 | 65.53 0.66 |
| Internal Test | External Test | |||
| Method | F1 (macro) | Acc | F1 (macro) | Acc |
| Global + Attention | 75.45 | 77.27 | 68.47 | 68.66 |
| Local + Attention | 79.04 | 81.81 | 71.16 | 71.76 |
| Common Attention | 78.65 | 79.54 | 73.06 | 73.61 |
| MultiToken Attention | 69.88 | 71.59 | 68.24 | 68.66 |
| Ours | 81.93 | 82.95 | 76.10 | 76.50 |
| Internal Test | External Test | |||
| Method | F1 (macro) | Acc | F1 (macro) | Acc |
| Fully-Connected Graph | 73.91 | 75.00 | 69.00 | 69.12 |
| Direct Fusion | 53.94 | 56.82 | 56.10 | 56.68 |
| Global Context | 58.26 | 61.36 | 47.02 | 52.07 |
| Ours (AOG) | 77.76 | 78.41 | 73.33 | 75.00 |
| Two Class | Three Class | |||
| Method | F1 (macro) | Acc | F1 (macro) | Acc |
| Radiologist 1 | 76.16 | 76.25 | 53.03 | 60.00 |
| Radiologist 2 | 74.98 | 75.00 | 53.63 | 60.00 |
| Ours | 76.25 | 76.25 | 59.71 | 62.50 |
| Group | Invasion-token attention |
| Correctly staged (T1–T2) | 0.058 |
| Correctly staged (T3–T4) | 0.105 |
| Over-staged (T1–T2 T3–T4) | 0.088 |
| Under-staged (T3–T4 T1–T2) | 0.058 |
| T-stage | Total | Positive invasion | Negative/absent |
| T1 | 151 | 20 (13.2%) | 131 (86.8%) |
| T2 | 116 | 25 (21.6%) | 91 (78.4%) |
| T3 | 99 | 52 (52.5%) | 47 (47.5%) |
| T4 | 73 | 50 (68.5%) | 23 (31.5%) |