OncoVision: Integrating Mammography and Clinical Data through Attention-Driven Multimodal AI for Enhanced Breast Cancer Diagnosis
Authors: Istiak Ahmed, Galib Ahmed, K. Shahriar Sanjid, Md. Tanzim Hossain, Md. Nishan Khan, Md. Misbah Khan, Md. Arifur Rahman, Sheikh Anisul Haque, +5 more
Organizations: Department Electrical and Computer Engineering, North South University, Dhaka, 1229, Bangladesh. · Big-Matrix Lab, North South University, Dhaka, 1229, Bangladesh. · Department of Mathematics and Physics, North South University, Dhaka, 1229, Bangladesh. · Department of Data Science, Friedrich-Alexander University, Erlangen, 91054, Germany. · Department of Oncology & Radiotherapy, Bangladesh Specialized Hospital, Shyamoli, Dhaka, 1207, Dhaka, Bangladesh. · Department of Transfusion Medicine, Bangladesh Specialized Hospital, Shyamoli, Dhaka, 1207, Dhaka, Bangladesh. · Department of Radiology & Imaging, Popular Medical College, 21 Shyamoli, Dhaka, 1205, Dhaka, Bangladesh. · Department of Transfusion Medicine, Khwaja Yunus Ali Medical College & Hospital, Enayetpur, Sirajganj, 6751, Rajshahi, Bangladesh. · Department of Radiology and Imaging, Bangladesh Medical University, Shahbag, Dhaka, 1000, Bangladesh. · Department of Oncology, University of Cambridge, Cambridge Biomedical Campus, Cambridge, CB2 0SP, United Kingdom.
OncoVision is a privileged-information training framework that uses mammography images and clinical features during training and performs inference from mammographic images alone. Employing an attention-based encoder-decoder backbone, it jointly segments four regions of interest (masses, calcifications, axillary findings, and breast tissue) with accuracy exceeding the nnU-Net baseline and predicts ten structured clinical features, including BI-RADS category. We developed two late-fusion strategies, Independent and Dependent, that integrate imaging, radiomic, and clinical information during training to improve diagnostic precision and potentially reduce inter-observer variability. Radiomic features extracted from predicted masks provide shape, intensity, and texture descriptors that complement the learned CNN representations. We evaluated OncoVision in a retrospective multi-reader study with six board-certified radiologists, assessing diagnostic confidence, reading time, and segmentation accuracy with and without AI assistance. In a paired reader-assistance evaluation, OncoVision was associated with higher diagnostic confidence for junior and senior radiologists, reduced reading time by up to 61%, and achieved segmentation accuracy comparable to or exceeding that of radiologists for mass lesions. We operationalized OncoVision as a secure web application, now deployed at a partner hospital, that generates structured reports with dual-confidence scoring and attention-weighted visualizations for real-time diagnostic support. The platform is designed for integration into clinical workflows, with the goal of supporting screening access in underprivileged regions. By combining accurate segmentation with clinical intuition, OncoVision advances AI-assisted mammographic interpretation, offering a scalable and accessible approach to earlier and more consistent image interpretation.
Figures & tables
Fig. 1 : Integrated workflow and multimodal architecture of OncoVision. (a) Integrated workflow from data acquisition and preprocessing to model training, evaluation, and web deployment. (b) Independent variant: raw clinical features and radiomic descriptors are concatenated with imaging-derived bottleneck features during training. (c) Dependent variant: clinical features are processed through a dedicated encoder before fusion with imaging and radiomic features. At inference, externally supplied clinical features are omitted, while radiomic features are generated internally from the predicted segmentation masks.
Fig. 2 : Detailed architecture of OncoVision. (a) Attention-gated U-Net segmentation backbone for four ROIs. (b) Independent variant: raw clinical features ( T ) and radiomic features ( R ) concatenated with imaging embeddings during training. (c) Dependent variant: clinical features processed through a dedicated encoder before fusion with imaging and radiomic features. At inference, externally supplied clinical features are omitted, while radiomic features are generated internally from the predicted segmentation masks and concatenated with the imaging representation before the corresponding inference head.
#
Configuration
Mass Dice ↑
Mass HD ↓
BI-RADS Macro F1 ↑
Clin. Acc ↑
0
nnU-Net + MLP (no fusion)
0.9353
10.53
0.5812
0.7184
1
nnU-Net + Independent
0.9353
10.53
0.6724
0.8031
2
nnU-Net + Dependent
0.9353
10.53
0.6887
0.8172
3
OncoVision + Dependent
0.9521
7.82
0.8104
0.8823
4
OncoVision + Independent
0.9521
7.82
0.7721
0.8604
5
OncoVision + Dependent + attention
0.9521
7.82
0.8241
0.8912
Table 1 : Clinical report prediction roadmap. Each row adds one component to the configuration above. Clinical accuracy is the mean accuracy across the ten structured features (macro-averaged F1 for multi-class tasks). BI-RADS Macro F1 corresponds to the ordinal BI-RADS category prediction. Mass Dice and HD remain constant across rows sharing the same segmentation backbone. Highlighted cells mark the best value per metric.
Backbone
Fusion
BI-RADS Macro F1 ↑
Clin. Acc ↑
nnU-Net
Independent
0.6724
0.8031
nnU-Net
Dependent
0.6887
0.8172
OncoVision
Independent
0.7721
0.8604
OncoVision
Dependent
0.8104
0.8823
Table 2 : Ablation on backbone and fusion strategy. OncoVision with Dependent fusion achieves the best clinical prediction performance. Bold indicates the best-performing configuration.
Configuration
BI-RADS Macro F1 ↑
Clin. Acc ↑
OncoVision + Dependent
0.8104
0.8823
+ Attention gates
0.8241
0.8912
Table 3 : Ablation on attention gates. Adding attention gates improves clinical prediction and BI-RADS Macro F1. Bold indicates the best-performing configuration.
Radiomic Strategy
BI-RADS Macro F1 ↑
Clin. Acc ↑
No radiomics
0.8241
0.8912
Concatenation
0.8552
0.9057
Table 4 : Ablation on radiomic integration. Concatenation of radiomic features improves clinical prediction and BI-RADS Macro F1. Bold indicates the best-performing configuration.
ROI
↑ IoU
↑ Dice
↑ Precision
↑ Sensitivity
↑ F1 Score
↑ Specificity
↓ HD (px)
↓ ASD (px)
↑ Boundary IoU
RVD (mean ± std)
RAVD (mean ± std)
nnU-Net
Mass
0.8785
0.9353
0.9423
0.9313
0.9361
0.9794
10.5253
1.3192
0.2973
−0.02±0.16
0.05±0.15
Axilla Findings
0.7413
0.8514
0.8728
0.8356
0.8531
0.9719
25.1785
4.1672
0.2992
−0.04±0.21
0.09±0.18
Breast Tissue
0.6138
0.7607
0.7854
0.7386
0.7615
0.9457
55.2629
2.0713
0.3222
+0.01±0.42
0.21±0.37
Calcification
0.6485
0.7868
0.8217
0.7561
0.7877
0.9811
95.0511
18.7892
0.2103
−0.22±0.39
0.25±0.36
OncoVision
Table 5 : Segmentation performance comparison between nnU-Net and OncoVision across four regions of interest (ROIs). Metrics include IoU, Dice, Precision, Sensitivity, F1, Specificity, Hausdorff Distance (HD), Average Surface Distance (ASD), Boundary IoU, and Relative Volume Difference (RVD/RAVD, reported as mean ± std). Arrows ( ↑ / ↓ ) indicate whether larger or smaller values are preferred. Bold indicates the best-performing model. OncoVision achieves consistent gains across all ROIs, with the largest improvements for calcification and mass segmentation.
Fig. 3 : Comparative segmentation and diagnostic profiling using OncoVision. (a–c) Side-by-side segmentation comparison of nnU-Net and OncoVision across mass (a), calcification (b), and axilla findings with breast tissue (c). Panels show, from left to right: original mammogram, ground truth, nnU-Net prediction, OncoVision prediction, nnU-Net error map, and OncoVision error map (true positives in green, false positives in red, false negatives in blue). Zoomed insets (3–4 × ) highlight improved boundary delineation and reduced false errors by OncoVision. (d) Clinical feature predictions by the Dependent variant for five patients. All predictions match ground truth except one BI-RADS downgrade (P2: 6 → 5). Data are from the test set of 754 mammograms (100 patients).
Fig. 4 : Comparative per-image segmentation and clinical feature confidence of nnU-Net and OncoVision. (a) Box plots of per-image IoU and DSC for calcification ( n=309 ), axilla findings ( n=369 ), breast tissue ( n=754 ), and mass ( n=328 ). Red dashed/solid lines: nnU-Net; blue dashed/solid lines: OncoVision. Wilcoxon signed-rank test; ∗∗∗P<0.001 ; ∗∗P<0.01 ; ∗P<0.05 . Large effect sizes for calcification (DSC: Cohen’s d=−0.877 ) and breast tissue (IoU: d=−0.857 ). (b) Box plots of model confidence across ten clinical features ( n=754 each) for the Dependent (green) and Independent (orange) variants after radiomic feature fusion. The Dependent variant showed significantly higher confidence across all ten features (Wilcoxon signed-rank test; p<0.0001 for nine features; p<0.01 for Calcification Distribution). Mean differences ranged from 0.0038 (Calcification Distribution) to 0.0316 (Mass Density), with large paired effect sizes for Mass Presence ( d=1.76 ), Mass Density ( d=1.04 ), and Axilla Findings ( d=1.02 ). These results show that radiomic feature fusion improved model confidence across both binary and multi-class clinical prediction tasks.
Clinical Feature
Metric
Dependent
Independent
Δ
Preferred
Binary Tasks (Yes/No)
Mass Presence
Accuracy
0.9610
0.9556
+0.0054
Dependent
Mass Calcification
Accuracy
0.9801
0.9829
−0.0028
Independent
Axillary Findings
Accuracy
0.9537
0.9322
+0.0215
Dependent
Calcification Presence
Accuracy
0.9356
0.9209
+0.0147
Dependent
Multi-Class Tasks (Macro-Averaged F1)
Table 6 : Clinical feature prediction performance of OncoVision’s two late-fusion variants with radiomic features. Binary tasks are reported as accuracy; multi-class tasks as macro-averaged F1. The Dependent variant outperforms the Independent variant on 9 of 10 features. Bold indicates the best-performing variant.
Fig. 5 : Paired reader-assistance evaluation of OncoVision as a decision-support tool. (a) Diagnostic confidence across ten clinical features for six radiologists (JR: junior; SR: senior; EXP: expert) and OncoVision. (b) Paired confidence with/without AI assistance. (c) Diagnostic time with/without AI. (d) Segmentation IoU vs. radiologists across four ROIs. (e) Batch segmentation time. Statistical details are provided in the Reader Study section.
Vision Transformers (ViT) have become the architecture of choice for many computer vision tasks, yet their performance in computer-aided diagnostics remains limited. Focusing on breast cancer detection from mammograms, we identify two main causes for this shortfall. First, medical images are high-resolution with small abnormalities, leading to an excessive number of tokens and making it difficult for the softmax-based attention to localize and attend to relevant regions. Second, medical image classification is inherently fine-grained, with low inter-class and high intra-class variability, where standard cross-entropy training is insufficient. To overcome these challenges, we propose a framework with three key components: (1) Region of interest (RoI) based token reduction using an object detection model to guide attention; (2) contrastive learning between selected RoI to enhance fine-grained discrimination through hard-negative based training; and (3) a DINOv2 pretrained ViT that captures localization-aware, fine-grained features instead of global CLIP representations. Experiments on public mammography datasets demonstrate that our method achieves superior performance over existing baselines, establishing its effectiveness and potential clinical utility for large-scale breast cancer screening. Our code is available for reproducibility here: https://aih-iitd.github.io/publications/attend-what-matters
Department of Computer Science and Engineering, IIT Delhi, New Delhi, India · Yardi School of AI, IIT Delhi, New Delhi, India · Department of Radiodiagnosis, PGIMER Chandigarh, Chandigarh, India
We present Aegis, a joint-embedding predictive architecture for breast cancer detection and density assessment in mammography. We train three Vision Transformer variants (Small/Base/Large) using self-supervised joint-embedding predictive architecture (JEPA) pre-training on 71,103 studies from 14 clinical sites, followed by supervised fine-tuning with progressive resolution scaling up to 2048x1536. On a curated 785-study test set, our largest model achieves area under the receiver operating characteristic curve (AUC) 0.949 for breast cancer triage with 93% sensitivity and 75% specificity at the optimal operating point. An ensemble combining our model with a U.S. Food and Drug Administration-cleared baseline further improves discrimination to 0.952 AUC. For breast density classification, the model achieves 0.953 AUC for binary (dense vs. non-dense) classification and 62.6% exact accuracy across four Breast Imaging Reporting and Data System (BI-RADS) categories, with 98.8% adjacent accuracy comparable to reported human inter-reader agreement. External validation on the public VinDr-Mammo dataset provides evidence of cross-population transfer under a different reference standard, with the largest model achieving 0.871 AUC for triage in a zero-shot setting.
Scott Chase Waggener, Sai Karthik Navuluru, Lakshman Tamil
Department of Electrical and Computer Engineering University of Texas at Dallas Richardson, TX
Mammography is an essential tool for breast cancer detection, with millions of examinations conducted annually. However, publicly available high-quality mammography datasets for AI development remain limited in both scale and annotation richness, particularly regarding pathological subtype coverage and structured diagnostic reasoning annotations. In this paper, we present MammoExpert, the first mammography dataset with Chain-of-Thought reasoning annotations across three diagnostic phases: (i) primal observation, (ii) factual assessment, and (iii) diagnostic synthesis. Comprising 2,379 mammography images covering 67 WHO-classified histopathology subtypes, each exam provides 42 radiographic features annotated by nine senior radiologists. We evaluate its performance on the breast lesion classification task, demonstrating superior accuracy and reasonability compared to existing classification models. Combining public dataset CBIS-DDSM with MammoExpert yields 7.1% classification accuracy improvement, while the training model to learn CoT reasoning achieves another 4% gain on the MammoExpert test set. Similar improvements are observed on INBreast and Vindr datasets, where the full approach yields accuracy gains of 6.9% and 6.7%, respectively. MammoExpert can serve as a benchmark for interpretable breast lesion diagnosis through explicit CoT reasoning.
Di Dai, Bo Liu, Youcheng Li +9
State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University Beijing, China · School of Computer Science and Engineer, Beijing University of Aeronautics and Astronautics Beijing, China · Center for Data Science, Peking University Beijing, China +5