cs.LGJul 22, 2026

Bayesian uncertainty estimation improves clinical decision making in medical AI agents

Authors: Frederik HaukePatrick WienholtChristiane KuhlDyke FerberJakob Nikolas KatherSven NebelungDaniel Truhn

Organizations: Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany. · Department of Medical Oncology, National Center for Tumor Diseases (NCT), Heidelberg University Hospital, Heidelberg, Germany. · Else Kröner Fresenius Center for Digital Health, TU Dresden, Dresden, Germany. · Department of Medicine I, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology,2026 Dresden, Germany.

Abstract

Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atypical cases. Here we show that Monte Carlo dropout, applied to a multi-task chest-radiograph classifier (eight thoracic findings, 137,593 training images), provides an epistemic uncertainty signal that tracks generalisation across training-set scales and flags confident yet error-prone predictions. Adding this signal to the point prediction raised error-detection AUROC from 0.74 to 0.77 (ΔΔAUROC +0.023, 95% CI [+0.014, +0.033]). In a controlled 2x2 factorial experiment, a clinical-decision-support agent exploited this uncertainty only when it was delivered as a binary error-risk flag rather than as raw scores, cutting confident misdiagnoses on unreliable findings from 8.5% to 2.7%. Epistemic uncertainty estimation thus carries decision-relevant information beyond point predictions, but its value for downstream agents depends on how it is communicated.

Explore similar work

CardsList