MedBayes-Lite: A Clinical Uncertainty Governance Layer for Risk-Aware Medical Decision Support
Authors: Elias Hossain, Md Mehedi Hasan Nipu, Maleeha Sheikh, Tasfia Nuzhat, Rajib Rana, Subash Neupane, Björn W. Schuller, Niloofar Yousefi
Organizations: College of Engineering and Computer Science, University of Central Florida, Orlando, FL 32816, USA · Department of Computer Science and Engineering, North South University, Dhaka 1229, Bangladesh · Department of Electrical and Computer Engineering, Purdue University Fort Wayne, Fort Wayne, IN 46805, USA · School of Mathematics, Physics and Computing, University of Southern Queensland, Springfield Central, QLD 4300, Australia · Meharry Medical College, Nashville, TN 37208, USA · CHI – Chair of Health Informatics, Technical University of Munich (TUM), Munich, Germany · GLAM – Group on Language, Audio, & Music, Imperial College London, London, UK
Abstract
Clinical language models often assign high confidence to incorrect predictions, particularly in high-severity and out-of-distribution cases. We present MedBayes-Lite, a retraining-free uncertainty governance layer for transformer-based clinical predictors. It combines Monte Carlo dropout, predictive calibration, and confidence-guided abstention to defer low-confidence predictions for human review, adding no trainable parameters. Evaluated on MedMCQA and MedQA-USMLE, MedBayes-Lite reduces expected calibration error by 0.23 to 0.33 and drives harmful overconfident errors (confident, incorrect, high-severity predictions) toward zero. Under domain shift from MedMCQA to MedQA-USMLE, it reduces confident high-severity errors from about 21% to near zero while roughly halving calibration drift. We also introduce the Clinical Uncertainty Score (CUS), which strongly correlates with harmful overconfidence (r approximately 0.88). Although the framework does not improve risk-coverage ranking, and temperature scaling or deep ensembles may provide advantages in calibration cost or risk ranking, MedBayes-Lite offers a practical calibration-and-abstention layer that reduces confident high-severity errors in clinical question-answering benchmarks.