Organizations: Lanzhou University Lanzhou, China · Shenzhen University Shenzhen, China · Southern Medical University Guangzhou, China · Central China Normal University Wuhan, China
Breast ultrasound diagnosis relies on clinically meaningful semantic concepts, yet most deep learning methods adopt end-to-end image-to-label paradigms that lack interpretability and robustness. While concept-based approaches offer a promising alternative, they often assume complete annotations or require multimodal inputs at inference, which significantly limits their real-world applicability. To tackle these issues, we propose Training-time Report-guided and Clinically Ordered Concept Editing (TRACE), a training-time report-guided framework that leverages structured radiology reports as privileged concept supervision while enabling image-only diagnosis at test time. TRACE refines image-derived concepts through a teacher-guided editing mechanism within a malignancy-aware ordered concept space. To address incomplete annotations, we introduce Strategic Concept Missing Training (SCMT) and train an image-only self-editor via edit distillation for autonomous concept refinement. Besides, we introduce BUSC, a concept-enriched benchmark linking images, labels, and structured attributes. Experiments across multiple datasets demonstrate that TRACE achieves superior performance and improved cross-domain robustness compared to existing methods.
Figures & tables
Figure 1. Overview of the structured concept representation used in BUSC-BUSBRA and performance comparison on the BUSC-BUSBRA dataset. The upper panel shows structured breast ultrasound concepts, including shape, margin, orientation, posterior features, echogenicity, and calcification, together with benign and malignant examples. The lower panel compares ResNet50, PCBM, and TRACE on AUC and accuracy. TRACE achieves an AUC of 91.6 percent and an accuracy of 85.9 percent, improving over PCBM by 1.9 and 1.8 percentage points, respectively.
Figure 2. Overview of TRACE. During training, structured reports act as privileged concept teachers to correct coarse image-derived concepts through residual concept editing. TRACE further improves robustness with hierarchical concept missing training and a clinically ordered risk space. At test time, the report branch is removed, and diagnosis is made from images alone via self-editing concept refinement. The TRACE framework consists of training and inference pathways. During training, breast ultrasound images produce coarse image-derived concepts, while structured reports provide privileged concept supervision for residual concept editing. Hierarchical concept missing training and a clinically ordered risk space improve robustness. During inference, the report branch is removed, and the image-only self-editor refines the concepts used for the final diagnosis.
Setting
Model
BUSC-BUSBRA
BUSC-BUSI647
AUC
Acc
F1
AUC
Acc
F1
Black-box
ResNet50 ( He et al., 2016 )
0.897 ± 0.015
0.826 ± 0.014
0.725 ± 0.027
0.945 ± 0.027
0.901 ± 0.026
0.838 ± 0.047
ViT-B/16 ( Dosovitskiy et al., 2020 )
0.876 ± 0.018
0.821 ± 0.025
0.711 ± 0.042
0.944 ± 0.021
0.886 ± 0.035
0.809 ± 0.068
Prototype
ProtoCaps ( Gallée et al., 2025 )
0.727 ± 0.037
0.718 ± 0.024
0.391 ± 0.075
0.803 ± 0.024
0.807 ± 0.014
0.658 ± 0.052
CBM
PCBM ( Yuksekgonul et al., 2023 )
0.886 ± 0.019
0.841 ± 0.018
0.738 ± 0.031
0.951 ± 0.022
0.901 ± 0.025
0.837 ± 0.047
MVP-CBM ( Wang et al., 2025 )
0.889 ± 0.020
0.830 ± 0.021
0.749 ± 0.027
0.931 ± 0.027
0.878 ± 0.045
0.806 ± 0.069
Table 1. Main comparison on BUSC-BUSBRA and BUSC-BUSI647. We compare TRACE with eight representative baselines. Best and second-best results in each column are highlighted in red and blue , respectively.
Model
Ardakani
BUS_UC
BrEaST
BUSC-BUSI647 ∗
AUC
Acc
AUC
Acc
AUC
Acc
AUC
Acc
ResNet50 ( He et al., 2016 )
0.857 ± 0.013
0.794 ± 0.025
0.525 ± 0.074
0.470 ± 0.053
0.508 ± 0.075
0.515 ± 0.125
0.873 ± 0.007
0.819 ± 0.016
ViT-B16 ( Dosovitskiy et al., 2020 )
0.844 ± 0.016
0.825 ± 0.017
0.657 ± 0.030
0.622 ± 0.026
0.771 ± 0.019
0.725 ± 0.029
0.847 ± 0.015
0.813 ± 0.012
ProtoCaps ( Gallée et al., 2025 )
0.756 ± 0.015
0.792 ± 0.016
0.538 ± 0.025
0.518 ± 0.024
0.671 ± 0.023
0.655 ± 0.008
0.764 ± 0.015
0.764 ± 0.010
PCBM ( Yuksekgonul et al., 2023 )
0.851 ± 0.015
0.808 ± 0.051
0.653 ± 0.042
0.589 ± 0.051
0.779 ± 0.111
0.691 ± 0.096
0.864 ± 0.010
0.826 ± 0.015
MVP-CBM ( Wang et al., 2025 )
0.863 ± 0.008
0.776 ± 0.027
0.647 ± 0.032
0.591 ± 0.036
0.769 ± 0.104
0.664 ± 0.118
0.871 ± 0.007
0.806 ± 0.016
Table 2. 5-fold cross-domain zero-shot performance ( mean ± std ). Best and second-best results are highlighted in red and blue .
Target
Editor
AUC
Acc
Ardakani
TRACE-MLP
0.862 ± 0.011
0.833 ± 0.026
TRACE-Att
0.862 ± 0.018
0.836 ± 0.046
TRACE-Gate
0.861 ± 0.014
0.853 ± 0.017
BUS_UC
TRACE-MLP
0.660 ± 0.011
0.614 ± 0.008
TRACE-Att
0.661 ± 0.027
0.604 ± 0.023
TRACE-Gate
0.645 ± 0.031
0.589 ± 0.018
Table 3. Cross-domain comparison of different TRACE editors. Models are trained on BUSBRA and tested on external datasets without target-domain fine-tuning. Best and second-best results are highlighted in red and blue .
Setting
Editor
BUSC-BUSBRA
BUSC-BUSI647
AUC
Acc
F1
AUC
Acc
F1
Full
Attention
0.924 ± 0.007
0.874 ± 0.016
0.794 ± 0.023
0.945 ± 0.023
0.893 ± 0.024
0.822 ± 0.071
Gating
0.921 ± 0.014
0.859 ± 0.012
0.783 ± 0.012
0.948 ± 0.016
0.902 ± 0.018
0.841 ± 0.036
MLP
0.916 ± 0.014
0.859 ± 0.013
0.776 ± 0.026
0.970 ± 0.030
0.938 ± 0.057
0.902 ± 0.083
Missing
Attention
0.924 ± 0.013
0.874 ± 0.019
0.802 ± 0.018
0.873 ± 0.023
0.857 ± 0.017
0.759 ± 0.025
Gating
0.926 ± 0.012
0.866 ± 0.017
0.784 ± 0.031
0.871 ± 0.018
0.865 ± 0.030
0.763 ± 0.042
Table 4. Core ablation of TRACE on BUSBRA and BUSC-BUSI647. We compare three editor designs under two training settings: Full (without concept missing) and Missing (with concept missing). Best and second-best results within each dataset and each setting are marked in red and blue , respectively.
Figure 3. t-SNE visualization of feature embeddings learned by VLG-CBM, Explicd, MVP-CBM, and TRACE. Four t-SNE plots visualize the feature embeddings learned by VLG-CBM, Explicd, MVP-CBM, and TRACE, respectively. Each plot shows the distribution and separation of samples from the diagnostic classes in the two-dimensional embedding space.
Editor
Ratio
AUC
Acc
F1
MLP
100%
0.9280 ± 0.0078
0.8834 ± 0.0173
0.8136 ± 0.0150
50%
0.9322 ± 0.0188
0.8754 ± 0.0126
0.7971 ± 0.0318
20%
0.9265 ± 0.0192
0.8733 ± 0.0178
0.7966 ± 0.0206
Attention
100%
0.9263 ± 0.0118
0.8807 ± 0.0120
0.8070 ± 0.0211
50%
0.9304 ± 0.0132
0.8770 ± 0.0165
0.8035 ± 0.0215
20%
0.9272 ± 0.0127
0.8770 ± 0.0098
0.8033 ± 0.0103
Table 5. Ablation on supervision ratio on BUSC-BUSBRA. We vary the proportion of training samples with structured concept supervision from 100% to 50% and 20%. Best and second-best results are highlighted in red and blue .
The Hong Kong University of Science and Technology · The Chinese University of Hong Kong · The Hong Kong University of Science and Technology (Guangzhou) +3