Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation
Authors: Nathan Le, Magdalini Paschali, Arogya Koirala, Andrew Johnston, Zhongnan Fang, David B. Larson, Akshay S. Chaudhari, Camila Gonzalez
Organizations: Department of Radiology, Stanford University, Stanford, CA, USA · Department of Biomedical Data Science, Stanford University, Stanford, CA, USA · Department of Anesthesia, Intensive Care Medicine, and Pain Medicine, Medical University of Vienna, Spitalgasse 23, 1090 Vienna, Austria
Radiology AI systems increasingly inform clinical decisions such as triage, follow-up imaging, and treatment planning. For these decisions to be made safely, model outputs must be well calibrated, meaning predicted probabilities accurately reflect true risk. Many standard techniques for improving calibration, such as MC Dropout and Deep Ensembles, require access to model parameters or retraining. However, proprietary clinical AI systems operate as black boxes, preventing access to the model's internals. To that end, we propose a model-agnostic framework for improving calibration of black-box models using clinically grounded test-time augmentation (TTA). Our framework applies geometric and physics-inspired 3D CT perturbations and learns probability-level aggregation strategies without access to model internals or the original training data. Across pulmonary embolism and intracranial hemorrhage detection tasks, DualTTA achieved the strongest overall calibration among TTA methods, reducing the Expected Calibration Error by 54% (0.239 -> 0.109) and 43% (0.051 -> 0.029), respectively, while requiring only input-output access. Additionally, DualTTA outperformed uncertainty estimation techniques that require access to model internals, such as Temperature Scaling, MC Dropout, and Deep Ensembles, in most calibration metrics. These results demonstrate that learned TTA aggregation can improve the calibration of clinical AI systems, providing a practical approach for improving the reliability of black-box medical AI.
Figures & tables
Figure 1: Proposed framework for improving black-box calibration. Left: a CT volume is perturbed by geometric and acquisition-based augmentations, and each perturbed volume is scored by the same frozen black-box model (identical black boxes), yielding one class probability estimate per augmentation (green/red bars). (A–F) The compared aggregation strategies combine these probabilities into a final, better-calibrated prediction (right side of each panel). Numbers denote aggregation weights, faded bars denote discarded predictions, and gray boxes group augmentations of the same type. In (F), DualTTA additionally blends the aggregated and baseline outputs via a learned coefficient ω .
PE
ICH
Method
ECE ↓
Brier ↓
ECE ↓
Brier ↓
No TTA
0.239 ± 0.017
0.231 ± 0.005
0.051 ± 0.025
0.081 ± 0.009
Temperature scaling
0.242 ± 0.014
0.231 ± 0.006
0.032 ± 0.005
0.076 ± 0.004
MC dropout
0.240 ± 0.017
0.231 ± 0.005
0.050 ± 0.025
0.080 ± 0.009
Deep ensembles
0.224 ± 0.013
0.219 ± 0.007
0.031 ± 0.012
0.069 ± 0.001
Equal TTA
0.161 ± 0.043
0.198 ± 0.013
0.058 ± 0.022
0.079 ± 0.007
Table 1: Expected Calibration Error (ECE) and Brier Scores for PE and ICH across different methods. Temperature scaling, MC dropout, and deep ensembles are included as non-black-box comparators. Values are mean ±std over three seeds.
Figure 2: Reliability diagrams for detecting pulmonary embolism (A) and intracranial hemorrhage (B) across TTA aggregation methods. Each curve shows the mean fraction of positive cases vs. the mean predicted probability across three model seeds. Shaded bands indicate the inter-seed variability (mean ± std across seeds). Adaptive (quantile-based) binning is used: bin edges are determined from each method’s pooled predictions across seeds such that each bin contains an approximately equal number of samples, avoiding unreliable estimates in sparsely populated probability regions. The better the calibration, the closer to the diagonal.
Figure 3: DualTTA ablation on ECE across augmentation types and severity levels for PE and ICH. (A) DualTTA trained with a single augmentation type (all 3 severity levels). (B) DualTTA trained with all 6 augmentation types at a single severity level. Dashed lines indicate DualTTA performance when trained with all 18 augmentations. Solid lines indicate ECE with No TTA . Error bars show mean ± std across three seeds. Using all augmentations yields the best calibration, particularly for PE, while individual types and severity levels show comparable performance for ICH.
Increasingly advanced data augmentation techniques have greatly aided clinical medical research, increasing data diversity and improving model generalization capabilities. Although most current basic models exhibit strong generalization abilities, image quality varies due to differences in equipment and operators. To address these challenges, we present SegTTA, a framework that improves medical image segmentation without model retraining by combining four augmentations (Gamma correction, Contrast enhancement, Gaussian blur, Gaussian noise) with weighted voting across multiple MedSAM2 checkpoints. Experiments demonstrate consistent improvements across three diverse datasets: healthy uterus segmentation, uterine myoma detection, and multi class hepatic structure segmentation. Ablation studies reveal that large organs benefit from intensity augmentations while small lesions require noise augmentations. The voting threshold controls the coverage precision trade off, enabling task specific optimization for different clinical requirements. Ultimately, on a multiclass hepatic vessel dataset, compared to MedSAM2 baselines, our method achieves an increase of 1.6 in mIoU and 1.9 in aIoU, along with a reduction of approximately 2.0 in HD95. Code will be available at https://github.com/AIGeeksGroup/SegTTA.
Yihong Yao, Chunlei Li, Canxuan Gang +4
AI Geeks · Qingdao Municipal Hospital · University of Chinese Academy of Sciences
Test-time prompt tuning (TPT) enables adaptation on a single test instance, achieving improved accuracy but often sacrificing calibration performance. Most existing calibration methods introduce additional regularization terms to promote dispersion across text embeddings and reduce calibration error, yet these methods often suffer from a drop in accuracy. Motivated by the well-calibrated nature of zero-shot predictions, we propose CoTS, a simple yet effective post-hoc calibration method that preserves accuracy. Specifically, CoTS applies temperature scaling to minimize the confidence gap between adapted and zero-shot predictions. To fully exploit the potential of multiple augmentations during adaptation, we introduce a weak-strong ensemble strategy that further boosts accuracy. We then apply CoTS to this ensemble, termed E-CoTS, to maintain its well-calibrated property. Extensive experiments on diverse datasets and backbones show that our approaches effectively mitigate miscalibration without compromising primary accuracy. For instance, E-CoTS reduces the average expected calibration error of TPT from 11.90% to 5.38% on ImageNet variants, while even increasing accuracy from 60.74% to 62.95%. Moreover, when integrated with existing calibration methods, E-CoTS usually enhances both accuracy and calibration simultaneously.
Yuwei Liang, Jian Liang, Dapeng Hu +2
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China · NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences, Beijing, China · Independent Researcher
Test-time prompt tuning (TPT) has emerged as a powerful paradigm, refining prompts for each test sample via entropy minimization (EM) over multiple augmented views. However, we identify a limitation in the standard EM-based adaptation: it inherently drives the model toward overconfident predictions disregarding sample-specific uncertainty, leading to significant calibration degradation. To address these limitations, we propose a new objective that replaces the conventional EM loss by aligning the original-view prediction with a target distribution derived from augmented views via cross-entropy, while adversarially incorporating the entropy of the target distribution to capture sample-specific uncertainty. Furthermore, to better construct this target distribution, we apply confidence-aware temperature scaling to each augmented-view prediction according to its confidence, sharpening confident predictions while softening uncertain ones. This formulation allows the model to increase confidence only when the target distribution is reliable, while preserving uncertainty when it reflects ambiguous or conflicting augmented-view predictions. Extensive experiments across diverse benchmarks demonstrate that our approach not only achieves state-of-the-art accuracy but also significantly improves model calibration.
Jungwon Choi, Hyeonseo Jang, Kibok Lee +1
Department of Computer Science and Engineering, Chung-Ang University · Department of Statistics and Data Science, Yonsei University