Organizations: Lanzhou University Lanzhou, China · Shenzhen University Shenzhen, China · Southern Medical University Guangzhou, China · Central China Normal University Wuhan, China
Breast ultrasound diagnosis relies on clinically meaningful semantic concepts, yet most deep learning methods adopt end-to-end image-to-label paradigms that lack interpretability and robustness. While concept-based approaches offer a promising alternative, they often assume complete annotations or require multimodal inputs at inference, which significantly limits their real-world applicability. To tackle these issues, we propose Training-time Report-guided and Clinically Ordered Concept Editing (TRACE), a training-time report-guided framework that leverages structured radiology reports as privileged concept supervision while enabling image-only diagnosis at test time. TRACE refines image-derived concepts through a teacher-guided editing mechanism within a malignancy-aware ordered concept space. To address incomplete annotations, we introduce Strategic Concept Missing Training (SCMT) and train an image-only self-editor via edit distillation for autonomous concept refinement. Besides, we introduce BUSC, a concept-enriched benchmark linking images, labels, and structured attributes. Experiments across multiple datasets demonstrate that TRACE achieves superior performance and improved cross-domain robustness compared to existing methods.
Figures & tables
Figure 1. Overview of the structured concept representation used in BUSC-BUSBRA and performance comparison on the BUSC-BUSBRA dataset. The upper panel shows structured breast ultrasound concepts, including shape, margin, orientation, posterior features, echogenicity, and calcification, together with benign and malignant examples. The lower panel compares ResNet50, PCBM, and TRACE on AUC and accuracy. TRACE achieves an AUC of 91.6 percent and an accuracy of 85.9 percent, improving over PCBM by 1.9 and 1.8 percentage points, respectively.
Figure 2. Overview of TRACE. During training, structured reports act as privileged concept teachers to correct coarse image-derived concepts through residual concept editing. TRACE further improves robustness with hierarchical concept missing training and a clinically ordered risk space. At test time, the report branch is removed, and diagnosis is made from images alone via self-editing concept refinement. The TRACE framework consists of training and inference pathways. During training, breast ultrasound images produce coarse image-derived concepts, while structured reports provide privileged concept supervision for residual concept editing. Hierarchical concept missing training and a clinically ordered risk space improve robustness. During inference, the report branch is removed, and the image-only self-editor refines the concepts used for the final diagnosis.
Setting
Model
BUSC-BUSBRA
BUSC-BUSI647
AUC
Acc
F1
AUC
Acc
F1
Black-box
ResNet50 ( He et al., 2016 )
0.897 ± 0.015
0.826 ± 0.014
0.725 ± 0.027
0.945 ± 0.027
0.901 ± 0.026
0.838 ± 0.047
ViT-B/16 ( Dosovitskiy et al., 2020 )
0.876 ± 0.018
0.821 ± 0.025
0.711 ± 0.042
0.944 ± 0.021
0.886 ± 0.035
0.809 ± 0.068
Prototype
ProtoCaps ( Gallée et al., 2025 )
0.727 ± 0.037
0.718 ± 0.024
0.391 ± 0.075
0.803 ± 0.024
0.807 ± 0.014
0.658 ± 0.052
CBM
PCBM ( Yuksekgonul et al., 2023 )
0.886 ± 0.019
0.841 ± 0.018
0.738 ± 0.031
0.951 ± 0.022
0.901 ± 0.025
0.837 ± 0.047
MVP-CBM ( Wang et al., 2025 )
0.889 ± 0.020
0.830 ± 0.021
0.749 ± 0.027
0.931 ± 0.027
0.878 ± 0.045
0.806 ± 0.069
Table 1. Main comparison on BUSC-BUSBRA and BUSC-BUSI647. We compare TRACE with eight representative baselines. Best and second-best results in each column are highlighted in red and blue , respectively.
Model
Ardakani
BUS_UC
BrEaST
BUSC-BUSI647 ∗
AUC
Acc
AUC
Acc
AUC
Acc
AUC
Acc
ResNet50 ( He et al., 2016 )
0.857 ± 0.013
0.794 ± 0.025
0.525 ± 0.074
0.470 ± 0.053
0.508 ± 0.075
0.515 ± 0.125
0.873 ± 0.007
0.819 ± 0.016
ViT-B16 ( Dosovitskiy et al., 2020 )
0.844 ± 0.016
0.825 ± 0.017
0.657 ± 0.030
0.622 ± 0.026
0.771 ± 0.019
0.725 ± 0.029
0.847 ± 0.015
0.813 ± 0.012
ProtoCaps ( Gallée et al., 2025 )
0.756 ± 0.015
0.792 ± 0.016
0.538 ± 0.025
0.518 ± 0.024
0.671 ± 0.023
0.655 ± 0.008
0.764 ± 0.015
0.764 ± 0.010
PCBM ( Yuksekgonul et al., 2023 )
0.851 ± 0.015
0.808 ± 0.051
0.653 ± 0.042
0.589 ± 0.051
0.779 ± 0.111
0.691 ± 0.096
0.864 ± 0.010
0.826 ± 0.015
MVP-CBM ( Wang et al., 2025 )
0.863 ± 0.008
0.776 ± 0.027
0.647 ± 0.032
0.591 ± 0.036
0.769 ± 0.104
0.664 ± 0.118
0.871 ± 0.007
0.806 ± 0.016
Table 2. 5-fold cross-domain zero-shot performance ( mean ± std ). Best and second-best results are highlighted in red and blue .
Target
Editor
AUC
Acc
Ardakani
TRACE-MLP
0.862 ± 0.011
0.833 ± 0.026
TRACE-Att
0.862 ± 0.018
0.836 ± 0.046
TRACE-Gate
0.861 ± 0.014
0.853 ± 0.017
BUS_UC
TRACE-MLP
0.660 ± 0.011
0.614 ± 0.008
TRACE-Att
0.661 ± 0.027
0.604 ± 0.023
TRACE-Gate
0.645 ± 0.031
0.589 ± 0.018
Table 3. Cross-domain comparison of different TRACE editors. Models are trained on BUSBRA and tested on external datasets without target-domain fine-tuning. Best and second-best results are highlighted in red and blue .
Setting
Editor
BUSC-BUSBRA
BUSC-BUSI647
AUC
Acc
F1
AUC
Acc
F1
Full
Attention
0.924 ± 0.007
0.874 ± 0.016
0.794 ± 0.023
0.945 ± 0.023
0.893 ± 0.024
0.822 ± 0.071
Gating
0.921 ± 0.014
0.859 ± 0.012
0.783 ± 0.012
0.948 ± 0.016
0.902 ± 0.018
0.841 ± 0.036
MLP
0.916 ± 0.014
0.859 ± 0.013
0.776 ± 0.026
0.970 ± 0.030
0.938 ± 0.057
0.902 ± 0.083
Missing
Attention
0.924 ± 0.013
0.874 ± 0.019
0.802 ± 0.018
0.873 ± 0.023
0.857 ± 0.017
0.759 ± 0.025
Gating
0.926 ± 0.012
0.866 ± 0.017
0.784 ± 0.031
0.871 ± 0.018
0.865 ± 0.030
0.763 ± 0.042
Table 4. Core ablation of TRACE on BUSBRA and BUSC-BUSI647. We compare three editor designs under two training settings: Full (without concept missing) and Missing (with concept missing). Best and second-best results within each dataset and each setting are marked in red and blue , respectively.
Figure 3. t-SNE visualization of feature embeddings learned by VLG-CBM, Explicd, MVP-CBM, and TRACE. Four t-SNE plots visualize the feature embeddings learned by VLG-CBM, Explicd, MVP-CBM, and TRACE, respectively. Each plot shows the distribution and separation of samples from the diagnostic classes in the two-dimensional embedding space.
Editor
Ratio
AUC
Acc
F1
MLP
100%
0.9280 ± 0.0078
0.8834 ± 0.0173
0.8136 ± 0.0150
50%
0.9322 ± 0.0188
0.8754 ± 0.0126
0.7971 ± 0.0318
20%
0.9265 ± 0.0192
0.8733 ± 0.0178
0.7966 ± 0.0206
Attention
100%
0.9263 ± 0.0118
0.8807 ± 0.0120
0.8070 ± 0.0211
50%
0.9304 ± 0.0132
0.8770 ± 0.0165
0.8035 ± 0.0215
20%
0.9272 ± 0.0127
0.8770 ± 0.0098
0.8033 ± 0.0103
Table 5. Ablation on supervision ratio on BUSC-BUSBRA. We vary the proportion of training samples with structured concept supervision from 100% to 50% and 20%. Best and second-best results are highlighted in red and blue .
Concept Bottleneck Models provide interpretable-by-design predictions by mediating diagnosis through human-understandable concepts, but in medical imaging, their trustworthiness is often limited by the quality and granularity of available supervision. In particular, predicted concept activations can be driven by irrelevant regions, leading to spatially unfaithful explanations. We study a data-centric spatially grounded Concept Bottleneck Model (SG-CBM) that leverages coarse lesion delineations as weak supervision to encourage anatomically plausible concept evidence. For breast ultrasound, we derive two clinically motivated zones from each lesion mask: (i) an in-lesion region of interest for morphology-related concepts and (ii) a posterior acoustic band for posterior phenomena. We train concept maps using a grouped spatial grounding objective and preserve semantic faithfulness with a linear bottleneck classifier. Across five-fold stratified group cross-validation, the proposed SG-CBM improves diagnostic AUROC and concept macro-AUROC while markedly increasing spatial alignment of concept evidence. We also perform a Train-corrupt/Test-clean annotation-quality stress test to quantify the impact of supervision quality on diagnosis and spatial faithfulness. Overall, the results underscore the need for data-quality-aware supervision design and systematic trustworthiness validation for deployable healthcare AI systems.
Moshiur Rahman Tonmoy, Dunren Che, Haitham Y. Adarbah +1
Department of Electrical Engineering and Computer Science, Texas A&M University–Kingsville (TAMUK), Kingsville, TX 78363, USA
Medical imaging models often operate as black boxes, limiting interpretability and systematic debugging. We introduce an easy-to-use, plug-and-play framework for concept-based interpretation and model refinement. By aligning a single-modality encoder to BioMedCLIP, we construct a Concept Bottleneck Model (CBM) that enables concept-level interventions. These interventions allow us to isolate causal versus spuriously correlated concepts, validate insights with domain experts, and generate counterfactual samples for targeted fine-tuning. We evaluate our framework on a Mayo Clinic ultrasound dataset and the CheXpert 5x200 chest X-ray dataset. Results demonstrate that concept intervention enables reliable model diagnosis while maintaining, and occasionally improving predictive performance via guided fine-tuning. Our findings highlight the practical value of this framework for controlled, interpretable refinement of clinical deep learning models.
Samrajya Thapa, Daniel J. Quest, Timothy L. Kline +3
Iowa State University, Ames IA, USA · Mayo Clinic, Rochester MN, USA
Medical imaging modalities such as ultrasound and X-ray are widely used in clinical practice, where diagnosis follows a structured, evidence-driven workflow aligned with standardized criteria. While multimodal large language models (MLLMs) show promise for automated medical report generation, most existing systems rely on end-to-end multimodal fusion without modeling clinically defined intermediate attributes, leading to limited grounding and interpretability. To address this issue, we propose CORAL (COncept-grounded ReAsoning with Localization), a multimodal framework that integrates spatial grounding and concept-level supervision into a unified reasoning process. CORAL employs a prompt-driven medical segmentation model to localize lesions and predicts multi-class clinical attributes through a Concept Bottleneck module. The resulting textual concept tokens are combined with mask-modulated visual features within an MLLM to enable structured report generation and diagnostic prediction. Experiments on BUS-CoT and IU X-ray datasets demonstrate consistent improvements in diagnostic accuracy, concept consistency, and report quality over strong general-purpose and medical MLLMs, indicating that concept-grounded reasoning better aligns generation with clinical decision processes.
Xinyue Xu, Hongbin Lin, Juangui Xu +6
The Hong Kong University of Science and Technology · The Chinese University of Hong Kong · The Hong Kong University of Science and Technology (Guangzhou) +3