Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
Authors: Zhuofan Xie, Zishan Lin, Jinliang Lin, Jie Qi, Shaohua Hong, Shuo Li
Organizations: School of Electronic Technology and Engineering, Xiamen University, Xiamen, China · School of Informatics, Xiamen University, Xiamen, China · School of Engineering, Case Western Reserve University, Cleveland, USA
Active Learning (AL) reduces annotation costs in medical imaging by selecting only the most informative samples for labeling, but suffers from cold-start when labeled data are scarce. Vision-Language Models (VLMs) address the cold-start problem via zero-shot predictions, yet their temperature-scaled softmax outputs treat text-image similarities as deterministic scores while ignoring inherent uncertainty, leading to overconfidence. This overconfidence misleads sample selection, wasting annotation budgets on uninformative cases. To overcome these limitations, the Similarity-as-Evidence (SaE) framework calibrates text-image similarities by introducing a Similarity Evidence Head (SEH), which reinterprets the similarity vector as evidence and parameterizes a Dirichlet distribution over labels. In contrast to a standard softmax that enforces confident predictions even under weak signals, the Dirichlet formulation explicitly quantifies lack of evidence (vacuity) and conflicting evidence (dissonance), thereby mitigating overconfidence caused by rigid softmax normalization. Building on this, SaE employs a dual-factor acquisition strategy: high-vacuity samples (e.g., rare diseases) are prioritized in early rounds to ensure coverage, while high-dissonance samples (e.g., ambiguous diagnoses) are prioritized later to refine boundaries, providing clinically interpretable selection rationales. Experiments on ten public medical imaging datasets with a 20% label budget show that SaE attains state-of-the-art macro-averaged accuracy of 82.57%. On the representative BTMRI dataset, SaE also achieves superior calibration, with a negative log-likelihood (NLL) of 0.425.
Figures & tables
Figure 1 : SaE addresses overconfidence in VLM-based AL by reframing text-image similarities as Dirichlet evidence. Left: Traditional VLM-based methods rely on temperature-scaled softmax, producing overconfident probabilities that mislead sample selection. Right: SaE calibrates similarities into evidence, enabling decomposition into vacuity (knowledge gaps) and dissonance (decision conflicts) for interpretable sample acquisition.
Figure 2 : Overview of the proposed SaE framework. A frozen VLM encodes images and PubMed-augmented class prompts to produce text-image similarities, which are mapped by a trainable SEH into Dirichlet evidence. We decompose this evidence into vacuity and dissonance, and use them to score unlabeled samples for AL. The inset shows one acquisition round: high-vacuity cases are prioritized early to cover unseen phenotypes, and high-dissonance boundary disputes are prioritized later for expert annotation.
Dataset
Organ(s)
K
# train/val/test
DermaMNIST [ 13 , 63 ]
Skin
7
7007/1003/2005
Kvasir [ 48 ]
Colon
8
2000/800/1200
RETINA [ 34 , 49 ]
Retina
4
2108/841/1268
LC25000 [ 8 ]
Lung, Colon
5
12500/5000/7500
CHMNIST [ 29 ]
Colorectal
8
2496/1000/1504
BTMRI [ 46 ]
Brain
4
2854/1141/1717
Table 1: Evaluation is conducted on ten diverse medical datasets covering nine distinct organs. K is the number of classes. Splits follow the official BiomedCoOp [ 35 ] benchmark.
Dataset
Random
PCB [ 3 ]
MedCoOp +Coreset [ 58 ]
MedCoOp +Entropy [ 22 ]
MedCoOp +BADGE [ 2 ]
BiomedCoOp [ 35 ]
Ours (SaE)
DermaMNIST
69.42 ± 0.5
71.07 ± 0.6
74.11 ± 0.7
74.56 ± 0.5
75.46 ± 0.6
62.59 ± 1.8
80.21 ± 0.4
Kvasir
71.10 ± 0.7
72.92 ± 0.9
80.83 ± 0.6
81.92 ± 0.5
81.42 ± 0.4
78.89 ± 1.2
88.58 ± 0.3
RETINA
51.48 ± 0.8
53.55 ± 0.5
62.78 ± 0.9
65.22 ± 0.6
66.88 ± 0.5
61.28 ± 1.1
75.22 ± 0.4
LC25000
93.92 ± 0.3
95.71 ± 0.2
96.93 ± 0.2
97.47 ± 0.2
97.25 ± 0.2
92.68 ± 0.6
99.23 ± 0.2
CHMNIST
69.75 ± 0.6
85.31 ± 0.5
77.99 ± 0.6
86.64 ± 0.4
87.70 ± 0.5
79.05 ± 2.2
91.03 ± 0.4
BTMRI
83.40 ± 0.4
85.50 ± 0.4
86.26 ± 0.3
89.92 ± 0.3
89.57 ± 0.4
83.30 ± 1.3
93.46 ± 0.2
Table 2 : SaE achieves state-of-the-art AL performance, consistently outperforming all baselines on ten datasets. Results show mean accuracy (%) and std. dev. across 5 seeds at a 20% budget. BiomedCoOp is a few-shot reference.
Variant
Macro avg. (%)
Random
68.01
+ Dual-factor score a
73.35
+ VLM (similarity for evidence) b
78.62
SaE: + SEH c
82.57
Table 3: Ablation study confirms the SEH is the most critical component for SaE’s performance gain. Macro average accuracy (%) is reported over 10 datasets at a 20% budget.
Dataset
t=1
t=2
t=3
t=4
t=5
t=5t=3
DermaMNIST
72.07
76.11
77.56
78.46
80.21
96.70
Kvasir
74.92
80.83
84.92
86.42
88.58
95.87
RETINA
54.55
64.78
68.22
69.88
75.22
90.69
LC25000
96.71
98.93
99.03
99.13
99.23
99.80
CHMNIST
74.81
82.99
86.64
89.70
91.03
95.18
BTMRI
86.50
88.26
92.92
93.02
93.46
99.42
Table 4: SaE exhibits rapid early-round convergence, confirming its ability to mitigate the cold-start problem. Top-1 accuracy (%) is shown after each round ( ρ=0.2 , T=5 ). The final column quantifies efficiency (ratio of t=3 to t=5 accuracy).
Figure 3 : SaE achieves the lowest and most stable training loss across all AL rounds. Training trajectories on BTMRI ( ρ=0.2 ) are compared. SaE avoids the high initial loss of BADGE and the volatility of BiomedCLIP+BADGE. Dashed red lines mark round transitions.
Figure 4 : SaE achieves superior calibration and mitigates VLM overconfidence. Reliability diagrams on BTMRI at a 20% label budget visualize empirical accuracy versus predicted confidence across 15 bins. Solid curves show accuracy, the dashed diagonal indicates perfect calibration, and blue bars represent the confidence distribution. SaE follows the diagonal most closely, BADGE [ 2 ] is mildly underconfident, and PCB [ 3 ] remains strongly overconfident.
Figure 5 : SaE provides superior sample efficiency, achieving higher accuracy with a 20% budget. The mean test accuracy on CHMNIST (5 seeds, 95% CI) is reported. SaE is shown at multiple budgets; baselines are at 20% for comparison.
Figure 6 : Architecture of the Similarity Evidence Head (SEH). SEH employs a dual-branch design to process image features x and similarity scores s . The image branch consists of two stacked MLP blocks (Linear-BN-ReLU-Dropout) to extract deep semantic cues, while the similarity branch uses a single MLP block. These features are concatenated and fused via a final linear layer followed by a Softplus activation to ensure the output evidence strength λ is strictly positive.
Method
Loss Formulation
Acc (%)
NLL
Entropy Only
LSEH=Lent
92.15
0.468
Difficulty Only
LSEH=Ldiff
92.90
0.441
SaE (Dual)
LSEH=Ldiff+βLent
93.46
0.425
Table 5 : Ablation of SEH loss components on BTMRI [ 46 ] . Ldiff is the difficulty matching term, and Lent is the entropy consistency term. The combination yields the best performance.
Regression Variant
Target Transform
Acc (%)
NLL
Log-form
−log(λ)
93.35
0.432
Inverse-form (Ours)
(λ+ϵ)−1
93.46
0.425
Table 6 : Comparison of regression target forms on BTMRI [ 46 ] . The inverse form (regressing 1/λ ) slightly outperforms the log form (regressing −logλ ) in both accuracy and calibration.
Dataset
Metric
β=0.1
β=0.3
β=0.5
β=0.7
β=1.0
BTMRI
Acc (%)
92.82
93.18
93.46
93.24
92.95
NLL
0.452
0.435
0.425
0.431
0.448
ECE
0.029
0.025
0.021
0.024
0.027
RETINA
Acc (%)
73.54
74.65
75.22
74.92
74.18
NLL
0.535
0.508
0.492
0.501
0.519
ECE
0.047
0.041
0.039
0.040
0.044
Table 7: Sensitivity of SaE to the loss weight β . Performance is robust across a wide range of β , with the optimal trade-off consistently observed at β=0.5 . Extreme values (0.1 or 1.0) tend to degrade both calibration (NLL/ECE) and accuracy.
ϵ
1×10−4
5×10−4
1×10−3
5×10−3
1×10−2
Acc (%)
93.41
93.44
93.46
93.38
93.25
Table 8: Sensitivity of SaE to the numerical stability parameter ϵ on BTMRI. SaE is largely insensitive to ϵ within a reasonable range ( 10−4 to 10−2 ), with peak performance at the default setting.
Figure 7 : Visual interpretability comparison on BTMRI [ 46 ] and DermaMNIST [ 13 , 63 ] . We visualize Grad-CAM [ 57 ] activation maps for the PCB [ 3 ] and our SaE. The red dashed contours indicate the ground-truth lesion regions. (a) Original input images. (b) PCB attention is often scattered, focusing on irrelevant background regions or failing to cover the entire lesion. (c) SaE generates highly focused and accurate attention maps that align closely with the pathological regions, confirming that our evidence-calibrated strategy successfully localizes clinical features.
Dataset
M=4
M=16
M=32
M=64
DermaMNIST
78.86
80.21
79.57
77.26
BTMRI
92.22
93.46
92.93
91.56
Table 9 : Impact of context length M on AL performance ( ρ=0.2 ). A moderate length ( M=16 ) achieves the best accuracy by balancing semantic capacity and overfitting. Performance drops at larger lengths ( M=32,64 ) due to overfitting on limited active learning data.
Schedule Strategy
wv(t)
wd(t)
Acc (%)
NLL
Dissonance-only
=0
=1
89.12
0.584
Vacuity-only
=1
=0
92.55
0.460
Static balanced
=0.5
=0.5
92.85
0.445
SaE (Dynamic)
1−T−1t−1
T−1t−1
93.46
0.425
Table 10 : Ablation of acquisition schedules on BTMRI [ 46 ] . The dynamic schedule outperforms all static variants. Dissonance-only fails due to cold-start instability, while Vacuity-only lacks late-stage refinement. Our dynamic strategy optimally bridges these two needs.
Department of Radiology, Weill Cornell Medicine, New York, NY, USA · Department of Electrical and Computer Engineering, Cornell University, Ithaca, NY, USA