Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
Organizations: Johns Hopkins University, School of Medicine · Massachusetts Institute of Technology · Microsoft Healthcare & Life Sciences
Abstract
As vision-language models (VLMs) are increasingly deployed in clinical decision support, more than accuracy is required: knowing when to trust their predictions is equally critical. Yet, a comprehensive and systematic investigation into the overconfidence of these models remains notably scarce in the medical domain. We address this gap through a comprehensive empirical study of confidence calibration in VLMs, spanning three model families (Qwen3-VL, InternVL3, LLaVA-NeXT), three model scales (2B--38B), and multiple confidence estimation prompting strategies, across three medical visual question answering (VQA) benchmarks. Our study yields three key findings: First, overconfidence persists across model families and is not resolved by scaling or prompting, such as chain-of-thought and verbalized confidence variants. Second, simple post-hoc calibration approaches, such as Platt scaling, reduce calibration error and consistently outperform the prompt-based strategy. Third, due to their (strict) monotonicity, these post-hoc calibration methods are inherently limited in improving the discriminative quality of predictions, leaving AUROC at the same level. Motivated by these findings, we investigate hallucination-aware calibration (HAC), which incorporates vision-grounded hallucination detection signals as complementary inputs to refine confidence estimates. We find that leveraging these hallucination signals improves both calibration and AUROC, with the largest gains on open-ended questions. On closed-ended questions, format-specific signals such as the logit margin show more informative. Overall, our findings suggest post-hoc calibration as standard practice for medical VLM deployment over raw confidence estimates, and highlight the practical usefulness of hallucination signals to enable more reliable use of VLMs in medical VQA.
Figures & tables
| Uncalibrated | Platt Scaling | Isotonic Regr. | HAC-Platt | HAC-Gate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Samp. | Verb. | Samp. | Verb. | Samp. | Verb. | Samp. | Verb. | Samp. | Verb. | ||
| Closed | Qwen3-VL-8B | 0.551 | 0.600 | 0.551 | 0.600 | 0.556 | 0.604 | 0.639 | 0.608 | 0.639 | 0.608 |
| InternVL3-8B | 0.663 | 0.522 | 0.663 | 0.522 | 0.660 | 0.522 | 0.677 | 0.545 | 0.678 | 0.545 | |
| LLaVA-NeXT-7B | 0.598 | 0.558 | 0.598 | 0.558 | 0.587 | 0.562 | 0.597 | 0.547 | 0.598 | 0.556 | |
| Open | Qwen3-VL-8B | 0.643 | 0.710 | 0.643 | 0.710 | 0.642 | 0.710 | 0.750 | 0.771 | 0.750 | 0.743 |
| InternVL3-8B | 0.764 | 0.614 | 0.764 | 0.614 | 0.763 | 0.620 | 0.789 | 0.742 | 0.793 | 0.736 | |
| Semantic Entropy | Predictive Entropy | Majority Vote | Distinct Answers | Eigen Score | UMPIRE | HAC-Platt | |||
|---|---|---|---|---|---|---|---|---|---|
| Samp. | Verb. | ||||||||
| Closed | Qwen3-VL-8B | 0.530 | 0.545 | 0.545 | 0.543 | 0.577 | 0.593 | 0.639 | 0.606 |
| InternVL3-8B | 0.601 | 0.661 | 0.661 | 0.644 | 0.509 | 0.606 | 0.677 | 0.542 | |
| LLaVA-NeXT-7B | 0.606 | 0.598 | 0.598 | 0.531 | 0.549 | 0.555 | 0.597 | 0.541 | |
| Open | Qwen3-VL-8B | 0.666 | 0.661 | 0.654 | 0.662 | 0.668 | 0.505 | 0.750 | 0.770 |
| InternVL3-8B | 0.751 | 0.770 | 0.750 | 0.770 | 0.413 | 0.684 | 0.789 | 0.747 | |
| Uncalibrated | Platt Scaling | Isotonic Regr. | HAC-Platt | HAC-Gate | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Samp. | Verb. | Samp. | Verb. | Samp. | Verb. | Samp. | Verb. | Samp. | Verb. | ||
| Closed | Qwen3-VL-8B | 0.214 | 0.180 | 0.108 | 0.109 | 0.096 | 0.095 | 0.101 | 0.108 | 0.101 | 0.115 |
| InternVL3-8B | 0.152 | 0.176 | 0.116 | 0.094 | 0.102 | 0.093 | 0.114 | 0.096 | 0.113 | 0.095 | |
| LLaVA-NeXT-7B | 0.165 | 0.371 | 0.127 | 0.123 | 0.131 | 0.111 | 0.124 | 0.140 | 0.131 | 0.163 | |
| Open | Qwen3-VL-8B | 0.400 | 0.363 | 0.098 | 0.095 | 0.094 | 0.087 | 0.086 | 0.075 | 0.086 | 0.075 |
| InternVL3-8B | 0.228 | 0.352 | 0.094 | 0.102 | 0.090 | 0.094 | 0.094 | 0.087 | 0.097 | 0.091 | |
| Closed | Open | |||||
| Method | #Tok/Gen | Total | #Tok/Gen | Total | ||
| Samp. | Base | 10–20 | 2.9 3.7 | 29.0 | 7.2 52.1 | 72.0 |
| CoT | 10–20 | 146.2 367.4 | 1462.0 | 144.5 408.0 | 1445.0 | |
| Hal. | VASE | 20 | 2.4 1.5 | 48.0 | 8.1 6.2 | 162.0 |
| Verbalized | Vanilla | 1 | 25.8 24.1 | 25.8 | 39.6 29.6 | 39.6 |
| Vanilla+CoT | 1 | 138.1 51.7 | 138.1 | 141.0 52.7 | 141.0 | |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | # Images | Open-ended | Closed-ended | Total QA |
|---|---|---|---|---|
| VQA-RAD | 203 | 200 | 251 | 451 |
| SLAKE-EN | 96 | 706 | 355 | 1061 |
| VQA-Med-2019 | 500 | 436 | 64 | 500 |
| Pooled | 799 | 1342 | 670 | 2012 |
| Term | Probability |
|---|---|
| Almost certain | 0.95 |
| Highly likely | 0.90 |
| Very good chance | 0.85 |
| Probable | 0.75 |
| Likely | 0.70 |
| Better than even | 0.60 |
| Qwen3-4B | Qwen3-30B | GPT-5 | Count | % of dis. |
| ✓ | ✓ | ✗ | ||
| ✗ | ✓ | ✗ | ||
| ✗ | ✓ | ✓ | ||
| ✗ | ✗ | ✓ | ||
| ✓ | ✗ | ✗ | ||
| Total | ||||
| Judge | TP | FP | FN | TN | Precision | Recall | F1 |
|---|---|---|---|---|---|---|---|
| Qwen3-4B (ours) | |||||||
| Qwen3-30B | |||||||
| GPT-5 |
| Type | Model | Acc. ( ) | Conf. | Gap ( ) | ECE ( ) | ACE ( ) | AUROC ( ) |
|---|---|---|---|---|---|---|---|
| Closed | Qwen3-VL-2B | 0.740 | 0.964 | 0.224 | 0.235 | 0.224 | 0.586 |
| Qwen3-VL-8B | 0.766 | 0.980 | 0.214 | 0.219 | 0.214 | 0.551 | |
| Qwen3-VL-32B | 0.779 | 0.991 | 0.212 | 0.219 | 0.212 | 0.521 | |
| InternVL3-2B | 0.696 | 0.857 | 0.162 | 0.181 | 0.194 | 0.693 | |
| InternVL3-8B | 0.782 | 0.910 | 0.128 | 0.151 | 0.152 | 0.663 | |
| InternVL3-38B | 0.775 | 0.842 | 0.067 | 0.106 | 0.122 | 0.662 |
| Type | Model | Acc. ( ) | Conf. | Gap ( ) | ECE ( ) | ACE ( ) | AUROC ( ) |
|---|---|---|---|---|---|---|---|
| Closed | Qwen3-VL-2B | 0.689 | 0.949 | 0.260 | 0.288 | 0.281 | 0.560 |
| Qwen3-VL-8B | 0.741 | 0.974 | 0.233 | 0.247 | 0.240 | 0.537 | |
| Qwen3-VL-32B | 0.765 | 0.975 | 0.210 | 0.230 | 0.213 | 0.551 | |
| InternVL3-2B | 0.633 | 0.837 | 0.204 | 0.232 | 0.242 | 0.687 | |
| InternVL3-8B | 0.757 | 0.887 | 0.130 | 0.188 | 0.194 | 0.616 | |
| InternVL3-38B | 0.745 | 0.798 | 0.053 | 0.124 | 0.171 | 0.666 |
| Type | Model | Acc. ( ) | Conf. | Gap ( ) | ECE ( ) | ACE ( ) | AUROC ( ) |
|---|---|---|---|---|---|---|---|
| Closed | Qwen3-VL-2B | 0.786 | 0.974 | 0.188 | 0.199 | 0.189 | 0.607 |
| Qwen3-VL-8B | 0.786 | 0.984 | 0.198 | 0.206 | 0.198 | 0.560 | |
| Qwen3-VL-32B | 0.803 | 1.000 | 0.197 | 0.197 | 0.197 | 0.500 | |
| InternVL3-2B | 0.738 | 0.870 | 0.132 | 0.175 | 0.181 | 0.686 | |
| InternVL3-8B | 0.800 | 0.927 | 0.127 | 0.158 | 0.152 | 0.674 | |
| InternVL3-38B | 0.825 | 0.871 | 0.046 | 0.107 | 0.120 | 0.695 |
| Type | Model | Acc. ( ) | Conf. | Gap ( ) | ECE ( ) | ACE ( ) | AUROC ( ) |
|---|---|---|---|---|---|---|---|
| Closed | Qwen3-VL-2B | 0.688 | 0.973 | 0.285 | 0.286 | 0.286 | 0.567 |
| Qwen3-VL-8B | 0.750 | 0.980 | 0.230 | 0.242 | 0.242 | 0.549 | |
| Qwen3-VL-32B | 0.703 | 1.000 | 0.297 | 0.296 | 0.296 | 0.500 | |
| InternVL3-2B | 0.703 | 0.866 | 0.162 | 0.210 | 0.199 | 0.739 | |
| InternVL3-8B | 0.781 | 0.905 | 0.124 | 0.163 | 0.179 | 0.762 | |
| InternVL3-38B | 0.609 | 0.848 | 0.239 | 0.338 | 0.312 | 0.463 |
| VQA-RAD | SLAKE | VQA-Med | Avg | ||||||||||||||||||||||
| Closed | Open | Closed | Open | Closed | Open | Closed | Open | ||||||||||||||||||
| ECE | ACE | AUC | ECE | ACE | AUC | ECE | ACE | AUC | ECE | ACE | AUC | ECE | ACE | AUC | ECE | ACE | AUC | ECE | ACE | AUC | ECE | ACE | AUC | ||
| Q-2B | Samp. (Base) | 0.288 | 0.281 | 0.560 | 0.430 | 0.401 | 0.666 | 0.199 | 0.189 | 0.607 | 0.298 | 0.280 | 0.690 | 0.286 | 0.286 | 0.567 | 0.461 | 0.450 | 0.704 | 0.235 | 0.224 | 0.586 | 0.356 | 0.345 | 0.695 |
| Samp. (CoT) | 0.244 | 0.221 | 0.637 | 0.321 | 0.324 | 0.782 | 0.233 | 0.231 | 0.671 | 0.275 | 0.262 | 0.716 | 0.291 | 0.285 | 0.782 | 0.338 | 0.331 | 0.710 | 0.230 | 0.225 | 0.668 | 0.292 | 0.286 | 0.734 | |
| Verb. (Vanilla) | 0.245 | 0.257 | 0.518 | 0.369 | 0.378 | 0.625 | 0.150 | 0.172 | 0.586 | 0.346 | 0.337 | 0.603 | 0.297 | 0.325 | 0.549 | 0.547 | 0.547 | 0.577 | 0.200 | 0.199 | 0.559 | 0.414 | 0.411 | 0.590 | |
| Verb. (Van.+CoT) | 0.315 | 0.320 | 0.641 | 0.514 | 0.514 | 0.730 | 0.290 | 0.293 | 0.570 | 0.431 | 0.425 | 0.644 | 0.315 | 0.288 | 0.516 | 0.442 | 0.442 | 0.720 | 0.302 | 0.299 | 0.596 | 0.447 | 0.446 | 0.684 | |
| Sampling (Closed) | Sampling (Open) | |||||
|---|---|---|---|---|---|---|
| Model | ||||||
| Qwen3-VL-2B | 4.47 0.48 | –0.49 0.08 | –3.12 0.45 | 2.51 0.10 | –0.49 0.04 | –1.81 0.08 |
| Qwen3-VL-8B | 4.62 0.47 | –0.60 0.06 | –3.21 0.45 | 3.34 0.17 | –0.94 0.02 | –2.37 0.17 |
| Qwen3-VL-32B | 3.26 0.63 | –0.48 0.05 | –1.90 0.62 | 3.51 0.64 | –1.00 0.03 | –2.44 0.63 |
| InternVL3-2B | 3.40 0.42 | –0.48 0.06 | –1.78 0.38 | 3.70 0.30 | –0.58 0.07 | –2.02 0.30 |
| InternVL3-8B | 3.73 0.55 | –0.13 0.08 | –2.00 0.48 | 2.65 0.25 | –0.70 0.07 | –1.15 0.26 |