cs.LGSep 24, 2026

Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation

Authors: Nathan Le, Magdalini Paschali, Arogya Koirala, Andrew Johnston, Zhongnan Fang, David B. Larson, Akshay S. Chaudhari, Camila Gonzalez

Organizations: Department of Radiology, Stanford University, Stanford, CA, USA · Department of Biomedical Data Science, Stanford University, Stanford, CA, USA · Department of Anesthesia, Intensive Care Medicine, and Pain Medicine, Medical University of Vienna, Spitalgasse 23, 1090 Vienna, Austria

Abstract

Radiology AI systems increasingly inform clinical decisions such as triage, follow-up imaging, and treatment planning. For these decisions to be made safely, model outputs must be well calibrated, meaning predicted probabilities accurately reflect true risk. Many standard techniques for improving calibration, such as MC Dropout and Deep Ensembles, require access to model parameters or retraining. However, proprietary clinical AI systems operate as black boxes, preventing access to the model's internals. To that end, we propose a model-agnostic framework for improving calibration of black-box models using clinically grounded test-time augmentation (TTA). Our framework applies geometric and physics-inspired 3D CT perturbations and learns probability-level aggregation strategies without access to model internals or the original training data. Across pulmonary embolism and intracranial hemorrhage detection tasks, DualTTA achieved the strongest overall calibration among TTA methods, reducing the Expected Calibration Error by 54% (0.239 -> 0.109) and 43% (0.051 -> 0.029), respectively, while requiring only input-output access. Additionally, DualTTA outperformed uncertainty estimation techniques that require access to model internals, such as Temperature Scaling, MC Dropout, and Deep Ensembles, in most calibration metrics. These results demonstrate that learned TTA aggregation can improve the calibration of clinical AI systems, providing a practical approach for improving the reliability of black-box medical AI.

Figures & tables

Explore similar work

CardsList
  1. SegTTA: Training-Free Test-Time Augmentation for Zero-Shot Medical Imaging Segmentation

    Apr 19, 2026Yihong Yao, Chunlei Li, Canxuan Gang +4Semi-Supervised Medical Image SegmentationData Augmentation

  2. Bridging the Confidence Gap: Temperature Scaling for Calibrating Test-Time Prompt Tuning

    Sep 15, 2026Yuwei Liang, Jian Liang, Dapeng Hu +2Soft Prompt TuningImagenet

  3. Rethinking the Test-Time Prompt Tuning Objective from the Perspective of Calibration

    Aug 31, 2026Jungwon Choi, Hyeonseo Jang, Kibok Lee +1Soft Prompt TuningCalibrated Uncertainty