As vision-language models (VLMs) are increasingly deployed in clinical decision support, more than accuracy is required: knowing when to trust their predictions is equally critical. Yet, a comprehensive and systematic investigation into the overconfidence of these models remains notably scarce in the medical domain. We address this gap through a comprehensive empirical study of confidence calibration in VLMs, spanning three model families (Qwen3-VL, InternVL3, LLaVA-NeXT), three model scales (2B--38B), and multiple confidence estimation prompting strategies, across three medical visual question answering (VQA) benchmarks. Our study yields three key findings: First, overconfidence persists across model families and is not resolved by scaling or prompting, such as chain-of-thought and verbalized confidence variants. Second, simple post-hoc calibration approaches, such as Platt scaling, reduce calibration error and consistently outperform the prompt-based strategy. Third, due to their (strict) monotonicity, these post-hoc calibration methods are inherently limited in improving the discriminative quality of predictions, leaving AUROC at the same level. Motivated by these findings, we investigate hallucination-aware calibration (HAC), which incorporates vision-grounded hallucination detection signals as complementary inputs to refine confidence estimates. We find that leveraging these hallucination signals improves both calibration and AUROC, with the largest gains on open-ended questions. On closed-ended questions, format-specific signals such as the logit margin show more informative. Overall, our findings suggest post-hoc calibration as standard practice for medical VLM deployment over raw confidence estimates, and highlight the practical usefulness of hallucination signals to enable more reliable use of VLMs in medical VQA.
Figures & tables
Figure 1
Figure 2: ACE across sampling-based and verbalized confidence extraction methods and their prompting variants (Respective prompts are in Appendix C.1 .), evaluated on the pooled medical VQA benchmarks and averaged across the 7/8B models. Except for Top-K variants in the verbalized approach, prompting variants do not consistently improve calibration. Full results, including per-model and per-dataset breakdowns with ECE and AUROC, are in Appendix E.2 .
Figure 3: Calibration errors (ECE and ACE) before and after post-hoc calibration (Platt scaling) for sampling-based and verbalized confidence on closed- and open-ended questions, evaluated on the pooled datasets. Full results are in Tables 16 and 17 (Appendix F.2 ).
Uncalibrated
Platt Scaling
Isotonic Regr.
HAC-Platt
HAC-Gate
Samp.
Verb.
Samp.
Verb.
Samp.
Verb.
Samp.
Verb.
Samp.
Verb.
Closed
Qwen3-VL-8B
0.551
0.600
0.551
0.600
0.556
0.604
0.639
0.608
0.639
0.608
InternVL3-8B
0.663
0.522
0.663
0.522
0.660
0.522
0.677
0.545
0.678
0.545
LLaVA-NeXT-7B
0.598
0.558
0.598
0.558
0.587
0.562
0.597
0.547
0.598
0.556
Open
Qwen3-VL-8B
0.643
0.710
0.643
0.710
0.642
0.710
0.750
0.771
0.750
0.743
InternVL3-8B
0.764
0.614
0.764
0.614
0.763
0.620
0.789
0.742
0.793
0.736
Table 2: AUROC ( ↑ ) comparison across post-hoc calibration methods on the pooled datasets. Platt scaling preserves the same AUROC as uncalibrated scores due to its strictly monotonic transformation. Isotonic regression also preserves nearly identical AUROC due to its non-decreasing mapping. In contrast, HAC improves AUROC by incorporating hallucination signals that rerank predictions. On average across all 8 models, HAC improves AUROC by 5.3 pp over uncalibrated baselines, with notably larger gains on open-ended questions (+7.3 pp), especially for verbalized confidence (+10.1 pp). Full results are in Table 15 (Appendix F.1 ).
Semantic Entropy
Predictive Entropy
Majority Vote
Distinct Answers
Eigen Score
UMPIRE
HAC-Platt
Samp.
Verb.
Closed
Qwen3-VL-8B
0.530
0.545
0.545
0.543
0.577
0.593
0.639
0.606
InternVL3-8B
0.601
0.661
0.661
0.644
0.509
0.606
0.677
0.542
LLaVA-NeXT-7B
0.606
0.598
0.598
0.531
0.549
0.555
0.597
0.541
Open
Qwen3-VL-8B
0.666
0.661
0.654
0.662
0.668
0.505
0.750
0.770
InternVL3-8B
0.751
0.770
0.750
0.770
0.413
0.684
0.789
0.747
Table 3: AUROC ( ↑ ) comparison of HAC against uncertainty baselines on the pooled datasets. Baseline uncertainty measures discriminate weakly on closed-ended questions, never exceeding AUROC 0.661 . In contrast, HAC attains the best overall AUROC, with the largest gains on open-ended questions (up to +10.2 pp over the strongest baseline for Qwen3-VL-8B).
Uncalibrated
Platt Scaling
Isotonic Regr.
HAC-Platt
HAC-Gate
Samp.
Verb.
Samp.
Verb.
Samp.
Verb.
Samp.
Verb.
Samp.
Verb.
Closed
Qwen3-VL-8B
0.214
0.180
0.108
0.109
0.096
0.095
0.101
0.108
0.101
0.115
InternVL3-8B
0.152
0.176
0.116
0.094
0.102
0.093
0.114
0.096
0.113
0.095
LLaVA-NeXT-7B
0.165
0.371
0.127
0.123
0.131
0.111
0.124
0.140
0.131
0.163
Open
Qwen3-VL-8B
0.400
0.363
0.098
0.095
0.094
0.087
0.086
0.075
0.086
0.075
InternVL3-8B
0.228
0.352
0.094
0.102
0.090
0.094
0.094
0.087
0.097
0.091
Table 4: ACE ( ↓ ) comparison across calibration methods on the pooled dataset. Each cell shows sampling or verbalized confidence. Overall, both HAC and standard calibration approaches achieve notable calibration gains comparable to uncalibrated baselines. Full results are in Table 16 (Appendix F.2 ).
Closed
Open
Method
N
#Tok/Gen
Total
#Tok/Gen
Total
Samp.
Base
10–20
2.9 ± 3.7
29.0
7.2 ± 52.1
72.0
CoT
10–20
146.2 ± 367.4
1462.0
144.5 ± 408.0
1445.0
Hal.
VASE
20
2.4 ± 1.5
48.0
8.1 ± 6.2
162.0
Verbalized
Vanilla
1
25.8 ± 24.1
25.8
39.6 ± 29.6
39.6
Vanilla+CoT
1
138.1 ± 51.7
138.1
141.0 ± 52.7
141.0
Table 5: Computational cost per question. N : the number of required generations per question. #Tok/Gen reports the mean ± std of output tokens per generation.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
# Images
Open-ended
Closed-ended
Total QA
VQA-RAD
203
200
251
451
SLAKE-EN
96
706
355
1061
VQA-Med-2019
500
436
64
500
Pooled
799
1342
670
2012
Appendix
Table 6: Statistics of the test splits of medical VQA datasets with open- and closed-ended (yes/no) questions.
Term
Probability
Almost certain
0.95
Highly likely
0.90
Very good chance
0.85
Probable
0.75
Likely
0.70
Better than even
0.60
Appendix
Table 7: Mapping from linguistic confidence terms to numerical probabilities.
Qwen3-4B
Qwen3-30B
GPT-5
Count
% of dis.
✓
✓
✗
27
49.1
✗
✓
✗
16
29.1
✗
✓
✓
8
14.5
✗
✗
✓
2
3.6
✓
✗
✗
2
3.6
Total
55
100.0
Appendix
Table 8: Inter-model disagreement patterns on the N=55 cases where the three judges do not unanimously agree. ✓/✗ denote accept/reject.
Judge
TP
FP
FN
TN
Precision
Recall
F1
Qwen3-4B (ours)
58
6
19
42
0.906
0.753
0.822
Qwen3-30B
74
12
3
36
0.860
0.961
0.908
GPT-5
43
2
34
46
0.956
0.558
0.705
Appendix
Table 9: Human validation of the LLM judges. Judge verdicts are compared against expert labels on N=125 annotated samples (all 70 unanimous-agreement cases plus all 55 inter-judge disagreements).
Figure 4: Sample size ablation for 2B models . ECE ( ↓ ), ACE ( ↓ ), and AUROC ( ↑ ) as a function of N′ , averaged across question types. Shaded bands show std over 1,000 simulations.
Figure 5: Sample size ablation for 7/8B models . Same setup as Figure 4 .
Figure 6: Sample size ablation for 30B+ models . Same setup as Figure 4 . Note: N′≤20 since the original sample size is N=20 .
Figure 7: ECE across different confidence extraction methods, including sampling and verbalized methods and their prompting variants. The evaluation is on the pooled medical VQA benchmarks and averaged across the 7/8B models.
Figure 8: AUROC across different confidence extraction methods, including sampling and verbalized methods and their prompting variants. The evaluation is on the pooled medical VQA benchmarks and averaged across the 7/8B models.
Type
Model
Acc. ( ↑ )
Conf.
Gap ( ↓ )
ECE ( ↓ )
ACE ( ↓ )
AUROC ( ↑ )
Closed
Qwen3-VL-2B
0.740
0.964
0.224
0.235
0.224
0.586
Qwen3-VL-8B
0.766
0.980
0.214
0.219
0.214
0.551
Qwen3-VL-32B
0.779
0.991
0.212
0.219
0.212
0.521
InternVL3-2B
0.696
0.857
0.162
0.181
0.194
0.693
InternVL3-8B
0.782
0.910
0.128
0.151
0.152
0.663
InternVL3-38B
0.775
0.842
0.067
0.106
0.122
0.662
Appendix
Table 10: Effect of model scale on accuracy, confidence, and calibration (sampling-based confidence, micro-averaged across VQA-RAD, SLAKE-EN, and VQA-Med-2019). Overconf. Gap = Mean Conf. − Accuracy. ↑ : higher is better, ↓ : lower is better.
Type
Model
Acc. ( ↑ )
Conf.
Gap ( ↓ )
ECE ( ↓ )
ACE ( ↓ )
AUROC ( ↑ )
Closed
Qwen3-VL-2B
0.689
0.949
0.260
0.288
0.281
0.560
Qwen3-VL-8B
0.741
0.974
0.233
0.247
0.240
0.537
Qwen3-VL-32B
0.765
0.975
0.210
0.230
0.213
0.551
InternVL3-2B
0.633
0.837
0.204
0.232
0.242
0.687
InternVL3-8B
0.757
0.887
0.130
0.188
0.194
0.616
InternVL3-38B
0.745
0.798
0.053
0.124
0.171
0.666
Appendix
Table 11: Effect of model scale on accuracy, confidence, and calibration (sampling-based confidence) on VQA-RAD. Overconf. Gap = Mean Conf. − Accuracy. ↑ : higher is better, ↓ : lower is better.
Type
Model
Acc. ( ↑ )
Conf.
Gap ( ↓ )
ECE ( ↓ )
ACE ( ↓ )
AUROC ( ↑ )
Closed
Qwen3-VL-2B
0.786
0.974
0.188
0.199
0.189
0.607
Qwen3-VL-8B
0.786
0.984
0.198
0.206
0.198
0.560
Qwen3-VL-32B
0.803
1.000
0.197
0.197
0.197
0.500
InternVL3-2B
0.738
0.870
0.132
0.175
0.181
0.686
InternVL3-8B
0.800
0.927
0.127
0.158
0.152
0.674
InternVL3-38B
0.825
0.871
0.046
0.107
0.120
0.695
Appendix
Table 12: Effect of model scale on accuracy, confidence, and calibration (sampling-based confidence) on SLAKE-EN. Overconf. Gap = Mean Conf. − Accuracy. ↑ : higher is better, ↓ : lower is better.
Type
Model
Acc. ( ↑ )
Conf.
Gap ( ↓ )
ECE ( ↓ )
ACE ( ↓ )
AUROC ( ↑ )
Closed
Qwen3-VL-2B
0.688
0.973
0.285
0.286
0.286
0.567
Qwen3-VL-8B
0.750
0.980
0.230
0.242
0.242
0.549
Qwen3-VL-32B
0.703
1.000
0.297
0.296
0.296
0.500
InternVL3-2B
0.703
0.866
0.162
0.210
0.199
0.739
InternVL3-8B
0.781
0.905
0.124
0.163
0.179
0.762
InternVL3-38B
0.609
0.848
0.239
0.338
0.312
0.463
Appendix
Table 13: Effect of model scale on accuracy, confidence, and calibration (sampling-based confidence) on VQA-Med-2019. Overconf. Gap = Mean Conf. − Accuracy. ↑ : higher is better, ↓ : lower is better.
VQA-RAD
SLAKE
VQA-Med
Avg
Closed
Open
Closed
Open
Closed
Open
Closed
Open
ECE ↓
ACE ↓
AUC ↑
ECE ↓
ACE ↓
AUC ↑
ECE ↓
ACE ↓
AUC ↑
ECE ↓
ACE ↓
AUC ↑
ECE ↓
ACE ↓
AUC ↑
ECE ↓
ACE ↓
AUC ↑
ECE ↓
ACE ↓
AUC ↑
ECE ↓
ACE ↓
AUC ↑
Q-2B
Samp. (Base)
0.288
0.281
0.560
0.430
0.401
0.666
0.199
0.189
0.607
0.298
0.280
0.690
0.286
0.286
0.567
0.461
0.450
0.704
0.235
0.224
0.586
0.356
0.345
0.695
Samp. (CoT)
0.244
0.221
0.637
0.321
0.324
0.782
0.233
0.231
0.671
0.275
0.262
0.716
0.291
0.285
0.782
0.338
0.331
0.710
0.230
0.225
0.668
0.292
0.286
0.734
Verb. (Vanilla)
0.245
0.257
0.518
0.369
0.378
0.625
0.150
0.172
0.586
0.346
0.337
0.603
0.297
0.325
0.549
0.547
0.547
0.577
0.200
0.199
0.559
0.414
0.411
0.590
Verb. (Van.+CoT)
0.315
0.320
0.641
0.514
0.514
0.730
0.290
0.293
0.570
0.431
0.425
0.644
0.315
0.288
0.516
0.442
0.442
0.720
0.302
0.299
0.596
0.447
0.446
0.684
Appendix
Table 14: Calibration quality of different confidence prompting strategies (uncalibrated) across all models. ECE, ACE, and AUROC are reported per dataset and question type, with Avg computed on the pooled dataset. Dashes indicate degenerate cases where the model achieved 0% accuracy, making calibration metrics undefined. No single prompting variant consistently outperforms the base sampling or verbalized methods across all settings, reconfirming the finding that prompting strategies alone do not reliably improve calibration. Bold : best, underline : second best per column within each model.
Table 15: Post-hoc calibration comparison: AUROC ( ↑ ) across calibration methods (5-fold CV). Each cell shows sampling / verbalized confidence. Bold = best per row.
Table 16: Post-hoc calibration comparison: ACE ( ↓ ) across calibration methods (5-fold CV). Each cell shows sampling or verbalized confidence. Bold = best per row.
Table 17: Post-hoc calibration comparison: ECE ( ↓ ) across calibration methods (5-fold CV). Each cell shows sampling / verbalized confidence. Bold = best per row.
Sampling (Closed)
Sampling (Open)
Model
a^
b^
d^
a^
b^
d^
Qwen3-VL-2B
4.47 ± 0.48
–0.49 ± 0.08
–3.12 ± 0.45
2.51 ± 0.10
–0.49 ± 0.04
–1.81 ± 0.08
Qwen3-VL-8B
4.62 ± 0.47
–0.60 ± 0.06
–3.21 ± 0.45
3.34 ± 0.17
–0.94 ± 0.02
–2.37 ± 0.17
Qwen3-VL-32B
3.26 ± 0.63
–0.48 ± 0.05
–1.90 ± 0.62
3.51 ± 0.64
–1.00 ± 0.03
–2.44 ± 0.63
InternVL3-2B
3.40 ± 0.42
–0.48 ± 0.06
–1.78 ± 0.38
3.70 ± 0.30
–0.58 ± 0.07
–2.02 ± 0.30
InternVL3-8B
3.73 ± 0.55
–0.13 ± 0.08
–2.00 ± 0.48
2.65 ± 0.25
–0.70 ± 0.07
–1.15 ± 0.26
Appendix
Table 18: Learned HAC-Platt parameters a^ , b^ , d^ for s(c,h)=σ(a^⋅c+b^⋅h+d^) , reported as mean ± std across 5-fold CV on the pooled dataset. a^≥0 and b^≤0 are satisfied in all cases.
Figure 9: Ablation on hallucination detection metrics used in HAC-Platt. The pooled dataset is used.
Figure 10: Cross-dataset ACE transfer for HAC-Platt. Each cell shows the ACE when calibration is fitted on one dataset (row) and evaluated on another (column). Bold = best per column. The “Uncal.” row shows raw confidence without calibration.
Figure 11: Cross-dataset ECE transfer for HAC-Platt. Same layout as Figure 10 .
Figure 12: Cross-dataset AUROC transfer for HAC-Platt. Same layout as Figure 11 .
Table 19: Auxiliary signal substitution on closed-ended questions on pooled datasets. HAC-Logit replaces VASE with the model’s internal yes/no logit margin −∣pyes−pno∣ , keeping all other components fixed.
Figure 13: AUROC ( ↑ ) versus calibration-set size Ncal , per model (rows) and question type (columns), pooled over the three datasets. Uncalibrated and Platt share the same AUROC, so their curves coincide and are flat in Ncal .
Figure 14: Adaptive Calibration Error (ACE, ↓ ) versus calibration-set size Ncal , per model (rows) and question type (columns), pooled over the three datasets. The dashed line is the uncalibrated baseline.