Epistemic Uncertainty-Aware Defect Detection for Quality Control in Medical Device Manufacturing
Authors: Raham A. Butt, Marco Romanelli, Roche C. de Guzman
Organizations: Department of Engineering, Hofstra University, Hempstead, NY 11549 USA · Department of Computer Science, Hofstra University, Hempstead, NY 11549 USA
Objective: We investigate whether accounting for epistemic uncertainty can improve the reliability of automated defect detection in medical device manufacturing. Methods: We consider a machine learning framework that operates on heterogeneous manufacturing and device-report data represented with Knowledge Graphs. To mitigate errors arising from uncertainty in the decision model, we analyze a principled rejection strategy to abstain from predictions whose estimated epistemic uncertainty exceeds a specified threshold. We evaluate the approach using standard synthetic benchmarks and real-world medical device report data. Results: The theoretical results establish the validity of the method characterizing the regimes under which it is expected to be effective. Empirically, the rejection strategy enables explicit control of coverage, that is, the proportion of samples for which the model issues predictions, while improving performance on the retained samples. On 266,170 real-world FDA MAUDE device reports, a 10% abstention rate reduces classification error by 48%, and more aggressive rejection (approximately 70% coverage) yields near-perfect accuracy on the retained samples. On standard synthetic manufacturing benchmarks, abstaining on 9% of the decisions, our approach reduces the risk up to 63% compared with the standard no-abstention approach. Conclusions: Abstaining from predictions with high epistemic uncertainty can provide a practical tool for controlling the reliability of machine learning-based defect detection, especially in high-stakes medical device manufacturing applications. Significance: Uncertainty-aware defect detection may support safer and more reliable quality assurance in medical device manufacturing by identifying cases that require additional inspection rather than issuing potentially harmful predictions.
Figures & tables
Fig. 1 : Illustrative example – The impact of principled abstention on the decision regions ( Figure 1(a) ), risk-coverage curve ( Figure 1(b) ), and classification margin histogram ( Figure 1(c) ) for a toy dataset with 2-dimensional features and 3 balanced classes. The model rejects observations concentrated near decision boundaries, where predictions are more uncertain. Lowering the coverage by 19% reduces the classification risk on retained predictions by 50% . The softmax-margin distributions further show that rejected observations generally have smaller margins than retained observations, with only limited overlap, indicating that the reservation mechanism effectively identifies uncertain predictions and trades coverage for lower selective risk. Further details are provided in Section III-A .
Fig. 2 : Synthetic benchmark data – Feature distributions under different levels of class posterior overlap. Figure 2(a) shows the setting for the benchmark introduced in [ 7 ] , while Figure 2(b) shows a more challenging setting with increased overlap between the class-conditional feature distributions, controlled by the target Bayes error parameter in Algorithm 1 .
Fig. 3 : Synthetic benchmark data – (Best) Accuracy bar plot over multiple seeds for the baseline benchmark, comparing the performance of a single model w/o and w/ abstention (target coverage ≥90% ). For the baseline setting, abstention increases mean accuracy from 95.10% to 98.17% ( +3.07 percentage points), reducing classification error by 62.65% . At higher overlap, it increases mean accuracy from 87.59% to 89.94% ( +2.35 percentage points), reducing classification error by 18.94% . The narrow error bars, which represent variability across multiple random seeds, indicate that performance remains consistent under different random initializations.
Fig. 4 : Synthetic benchmark data – Risk as a function of coverage. Results are shown for the baseline setting ( Figure 4(a) ) and the more challenging setting with increased class posterior overlap ( Figure 4(b) ). In both cases, the risk decreases as the coverage decreases, indicating that the abstention mechanism preferentially rejects predictions that are more likely to be misclassified. Results over several random seeds are shown, with the shaded area indicating the standard deviation, showing that the results are consistent across different independent runs.
Fig. 5 : Medical device data – Performance across different values of k and coverage.
Fig. 6 : Medical device data – Performance comparison between no abstention and abstention. Results are reported over several random seeds, showing narrow error bars indicating consistent results across multiple independent runs. The abstention mechanism allows for an increase in average accuracy, from 94.68% to 97.23% , an improvement of 2.55 percentage points, corresponding to a 47.93% relative reduction in classification error with a target coverage of 90% , further demonstrating its effectiveness.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Fig. 7 : Synthetic benchmark data – Ablation study over the variable k for the higher posterior overlap case for the benchmark in Section IV-A .
Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atypical cases. Here we show that Monte Carlo dropout, applied to a multi-task chest-radiograph classifier (eight thoracic findings, 137,593 training images), provides an epistemic uncertainty signal that tracks generalisation across training-set scales and flags confident yet error-prone predictions. Adding this signal to the point prediction raised error-detection AUROC from 0.74 to 0.77 (ΔAUROC +0.023, 95% CI [+0.014, +0.033]). In a controlled 2x2 factorial experiment, a clinical-decision-support agent exploited this uncertainty only when it was delivered as a binary error-risk flag rather than as raw scores, cutting confident misdiagnoses on unreliable findings from 8.5% to 2.7%. Epistemic uncertainty estimation thus carries decision-relevant information beyond point predictions, but its value for downstream agents depends on how it is communicated.
Frederik Hauke, Patrick Wienholt, Christiane Kuhl +4
Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany. · Department of Medical Oncology, National Center for Tumor Diseases (NCT), Heidelberg University Hospital, Heidelberg, Germany. · Else Kröner Fresenius Center for Digital Health, TU Dresden, Dresden, Germany. +1
Machine learning in high-stakes domains such as healthcare requires not only strong predictive performance but also reliable uncertainty quantification (UQ) to support human oversight. Multi-label text classification (MLTC) is a central task in this domain, yet remains challenging due to label imbalances, dependencies, and combinatorial complexity. Existing MLTC benchmarks are increasingly saturated and may be affected by training data contamination, making it difficult to distinguish genuine reasoning capabilities from memorization. We introduce MADE, a living MLTC benchmark derived from {m}edical device {ad}verse {e}vent reports and continuously updated with newly published reports to prevent contamination. MADE features a long-tailed distribution of hierarchical labels and enables reproducible evaluation with strict temporal splits. We establish baselines across more than 20 encoder- and decoder-only models under fine-tuning and few-shot settings (instruction-tuned/reasoning variants, local/API-accessible). We systematically assess entropy-/consistency-based and self-verbalized UQ methods. Results show clear trade-offs: smaller discriminatively fine-tuned decoders achieve the strongest head-to-tail accuracy while maintaining competitive UQ; generative fine-tuning delivers the most reliable UQ; large reasoning models improve performance on rare labels yet exhibit surprisingly weak UQ; and self-verbalized confidence is not a reliable proxy for uncertainty. Our work is publicly available at https://hhi.fraunhofer.de/aml-demonstrator/made-benchmark.
Raunak Agarwal, Markus Wenzel, Simon Baur +3
Department of Artificial Intelligence Fraunhofer Heinrich Hertz Institute Berlin, 10587, Germany
Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable risk, regret, coverage surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.
Jakub Paplhám, Willem Waegeman, Eyke Hüllermeier +1
Czech Technical University in Prague · Ghent University · LMU Munich, MCML, DFKI