Medical imaging models often operate as black boxes, limiting interpretability and systematic debugging. We introduce an easy-to-use, plug-and-play framework for concept-based interpretation and model refinement. By aligning a single-modality encoder to BioMedCLIP, we construct a Concept Bottleneck Model (CBM) that enables concept-level interventions. These interventions allow us to isolate causal versus spuriously correlated concepts, validate insights with domain experts, and generate counterfactual samples for targeted fine-tuning. We evaluate our framework on a Mayo Clinic ultrasound dataset and the CheXpert 5x200 chest X-ray dataset. Results demonstrate that concept intervention enables reliable model diagnosis while maintaining, and occasionally improving predictive performance via guided fine-tuning. Our findings highlight the practical value of this framework for controlled, interpretable refinement of clinical deep learning models.
Figures & tables
Figure 1 : Overview of the proposed framework. A single-modality medical imaging model is linearly aligned to the BioMedCLIP embedding space to enable concept-based interpretability via a concept bottleneck model (CBM). Concept interventions on held-out data identify causal versus spurious concepts through prediction flips. These interventions are then used to augment data and fine-tune the CBM classification head, improving predictive performance.
Model Configuration
Test Set
Accuracy
F1
AUROC
Mayo Ensemble (Baseline)
Test Set 1
0.890
0.880
0.920
Test Set 2
0.730
0.740
0.840
CBM (Pre-Intervention)
Test Set 1
0.850
0.850
0.916
Test Set 2
0.740
0.740
0.839
CBM (Post-Intervention)
Test Set 1
0.850
0.850
0.917
Test Set 2
0.753
0.749
0.839
Table 1 : Performance evaluation on held-out Mayo Clinic test splits comparing the baseline ensemble, pre-intervention, and post-intervention fine-tuned CBM.
Pathology Class
TorchXrayVision Baseline
CBM (Pre-Intervention)
CBM (Post-Intervention)
Prec
Rec
F1
AUC
Prec
Rec
F1
AUC
Prec
Rec
F1
AUC
Cardiomegaly
0.63
0.68
0.65
0.86
0.73
0.58
0.65
0.85
0.73
0.58
0.65
0.85
Consolidation
0.50
0.05
0.09
0.71
0.39
0.71
0.51
0.69
0.43
0.71
0.54
0.75
Edema
0.43
0.21
0.28
0.70
0.48
0.45
0.47
0.83
0.48
0.50
0.49
0.80
Effusion
0.35
0.66
0.46
0.77
0.73
0.61
0.67
0.89
0.73
0.61
0.67
0.90
Atelectasis
0.43
0.63
0.51
0.76
0.91
0.50
0.65
0.84
0.92
0.55
0.69
0.90
Table 2 : CheXpert per-label performance comparison between TorchXrayVision baseline and our CBM framework before and after intervention-based fine-tuning.
Model
Accuracy
Macro F1
AUROC
Interpretability
CLIP Zero-shot
33.9
27.4
62.3
No
Aligned Model (Zero-shot)
33.3
24.5
64.1
Limited
X-ray Vision (Linear Probe)
53.7
52.0
82.1
No
CBM (Pre-intervention)
57.0
59.0
82.0
Yes
CBM (Post-intervention)
59.0
61.0
84.0
Yes
Table 3 : Overall performance comparison across models on CheXpert 5×200.
Figure 2 : Interpretation and intervention case studies across ultrasound (Mayo Clinic) and chest X-ray (CheXpert) modalities
Deep learning has revolutionized medical image analysis, delivering exceptional diagnostic accuracy across diverse applications. Yet, the lack of interpretability in its decision-making hinders clinical adoption, particularly in high-stakes medical contexts where transparency is paramount for trustworthiness. For example, in Placenta Accreta Spectrum (PAS), subtle cues in ultrasound imaging challenge reliable diagnosis, rendering black-box models untrustworthy for accurate scoring. To address this, Concept Bottleneck Models (CBMs) offer a promising avenue by embedding clinically meaningful intermediate concepts into the diagnosis pipeline, enabling clinicians to scrutinize and refine model outputs. However, conventional CBMs falter in capturing complex inter-concept dependencies and demand costly, expert-driven concept annotations, limiting their scalability. This study introduces a novel semi-supervised CBM framework designed for medical imaging, which leverages dual-level hypergraph learning to model high-order concept dependencies and generate domain-adaptive pseudo-labels. Our approach achieves superior interpretability and performance by integrating a concept-level hypergraph for enhanced reasoning and an image-level hypergraph for robust pseudo-label generation. Experiments on a newly annotated PAS ultrasound dataset and a breast ultrasound public dataset demonstrate the effectiveness of the proposed concept label-efficient interpretable framework. Its universality is further validated on the dermoscopic image dataset SkinCon. The code is available at https://github.com/scott-yjyang/HyperCBM.
Concept Bottleneck Models (CBMs) offer interpretable alternatives to black-box predictors by introducing human-relatable concepts before the final output. However, existing CBMs struggle to verify whether predicted concepts correspond to the correct visual evidence, limiting their reliability. We propose a fine-grained CBM framework that grounds each concept in localized visual evidence, enabling direct inspection of where and how concepts are encoded. This design allows users to interpret predictions and verify that the model learns intended concepts rather than spurious correlations. Experiments on medical imaging benchmarks show that our learned concept space is information-complete and achieves predictive performance comparable to standard CBMs, while substantially improving transparency. Unlike post-hoc attribution methods, our framework validates both the presence and correctness of concept representations, bridging interpretability with verifiability. Our approach enhances the trustworthiness of CBMs and establishes a principled mechanism for human-model interaction at the concept level, paving the way toward more reliable and clinically actionable concept-based learning systems.
Yingying Fang, Haijie Xu, Shuang Wu +2
Bioengineering Department and Imperial-X, Imperial College London, London, UK · Thoughtworks AI Labs, Singapore
Concept bottleneck models (CBMs) can improve the transparency of cancer image diagnostic prediction by expressing predictions through radiological concepts. However, their dependence on instance-level concept annotations limits practical applicability. We propose a prior-guided hybrid CBM that integrates limited concept annotations, class-conditional concept distribution matching on unannotated patients, and prior initialization of the concept-to-diagnosis head. We evaluate the method on CBIS-DDSM mammographic masses and calcifications and LIDC-IDRI pulmonary nodules across 0-100% concept annotation. In the clinically relevant 0-20% annotation regime, the hybrid CBM consistently improves mean concept AUC over a matched standard CBM, while maintaining diagnostic performance close to black-box models. At 10% annotation specifically, concept AUC increases from 0.619 to 0.741 for masses, from 0.650 to 0.787 for calcifications, and from 0.597 to 0.642 for pulmonary nodules. Ablation experiments identify prior initialization as the main component contributing to improved concept detection, likely by stabilizing the concept-to-diagnosis head. Zero-shot VLMs remain insufficient for reliable fine-grained tumor-level concept prediction. These findings suggest that structured priors can substantially reduce the annotation burden of interpretable cancer imaging models.
Baoqiang Ma, Kenneth Gilhuijs
Image Sciences Institute, University Medical Center Utrecht, Utrecht, The Netherlands