Medical imaging models often operate as black boxes, limiting interpretability and systematic debugging. We introduce an easy-to-use, plug-and-play framework for concept-based interpretation and model refinement. By aligning a single-modality encoder to BioMedCLIP, we construct a Concept Bottleneck Model (CBM) that enables concept-level interventions. These interventions allow us to isolate causal versus spuriously correlated concepts, validate insights with domain experts, and generate counterfactual samples for targeted fine-tuning. We evaluate our framework on a Mayo Clinic ultrasound dataset and the CheXpert 5x200 chest X-ray dataset. Results demonstrate that concept intervention enables reliable model diagnosis while maintaining, and occasionally improving predictive performance via guided fine-tuning. Our findings highlight the practical value of this framework for controlled, interpretable refinement of clinical deep learning models.
Figures & tables
Figure 1 : Overview of the proposed framework. A single-modality medical imaging model is linearly aligned to the BioMedCLIP embedding space to enable concept-based interpretability via a concept bottleneck model (CBM). Concept interventions on held-out data identify causal versus spurious concepts through prediction flips. These interventions are then used to augment data and fine-tune the CBM classification head, improving predictive performance.
Model Configuration
Test Set
Accuracy
F1
AUROC
Mayo Ensemble (Baseline)
Test Set 1
0.890
0.880
0.920
Test Set 2
0.730
0.740
0.840
CBM (Pre-Intervention)
Test Set 1
0.850
0.850
0.916
Test Set 2
0.740
0.740
0.839
CBM (Post-Intervention)
Test Set 1
0.850
0.850
0.917
Test Set 2
0.753
0.749
0.839
Table 1 : Performance evaluation on held-out Mayo Clinic test splits comparing the baseline ensemble, pre-intervention, and post-intervention fine-tuned CBM.
Pathology Class
TorchXrayVision Baseline
CBM (Pre-Intervention)
CBM (Post-Intervention)
Prec
Rec
F1
AUC
Prec
Rec
F1
AUC
Prec
Rec
F1
AUC
Cardiomegaly
0.63
0.68
0.65
0.86
0.73
0.58
0.65
0.85
0.73
0.58
0.65
0.85
Consolidation
0.50
0.05
0.09
0.71
0.39
0.71
0.51
0.69
0.43
0.71
0.54
0.75
Edema
0.43
0.21
0.28
0.70
0.48
0.45
0.47
0.83
0.48
0.50
0.49
0.80
Effusion
0.35
0.66
0.46
0.77
0.73
0.61
0.67
0.89
0.73
0.61
0.67
0.90
Atelectasis
0.43
0.63
0.51
0.76
0.91
0.50
0.65
0.84
0.92
0.55
0.69
0.90
Table 2 : CheXpert per-label performance comparison between TorchXrayVision baseline and our CBM framework before and after intervention-based fine-tuning.
Model
Accuracy
Macro F1
AUROC
Interpretability
CLIP Zero-shot
33.9
27.4
62.3
No
Aligned Model (Zero-shot)
33.3
24.5
64.1
Limited
X-ray Vision (Linear Probe)
53.7
52.0
82.1
No
CBM (Pre-intervention)
57.0
59.0
82.0
Yes
CBM (Post-intervention)
59.0
61.0
84.0
Yes
Table 3 : Overall performance comparison across models on CheXpert 5×200.
Figure 2 : Interpretation and intervention case studies across ultrasound (Mayo Clinic) and chest X-ray (CheXpert) modalities