cs.CVOct 1, 2026

Uncertainty-Guided Handshake: Efficient Human-in-the-Loop Refinement for Surgical-Grade Glioma Segmentation

Authors: Samuel Hart, Ahmad Yahya, Ahmed Karam Eldaly

Organizations: Department of Computer Science, University of Exeter, Exeter, EX4 4QF, United Kingdom. · Department of Nuclear Engineering, Faculty of Engineering, King Abdulaziz University, Jeddah, Saudi Arabia. · UCL Hawkes Institute, University College London, Gower St., London, WC1E 6AE, United Kingdom.

Abstract

While state-of-the-art automated models for medical image segmentation achieve high mean performance, they frequently suffer from localized, catastrophic failures that preclude safe clinical deployment, particularly in neuro-oncology. Interactive segmentation frameworks mitigate this by incorporating human oversight, but traditionally impose prohibitive cognitive and temporal workloads by requiring clinicians to manually search for errors. In this project, we present an efficient, Hybrid Structural-Aleatoric Human-in-the-Loop framework for glioma segmentation that bridges the gap between automated baseline performance and surgical-grade precision, achieving sub-2.0 mm HD95 on curated benchmarks while providing safety-net routing for structural failures across real-world clinical data. By extracting voxel-wise Test-Time Augmentation (TTA) uncertainty and applying hierarchical topological filtering, our method proactively isolates high-risk structural anomalies. We comprehensively evaluated our approach on a challenging out-of-distribution clinical stress-test cohort (N = 362). Operating under a simulated Human Oracle, the framework improved the Whole Tumor (WT) Dice score from 0.891 to 0.914 and reduced the 95th percentile Hausdorff Distance (HD95) from 5.82 mm to 4.76 mm. Critically for surgical safety, the system rescued severe boundary failures in the Tumor Core, reducing mean HD95 from 17.96 mm to 14.83 mm (improving absolute TC Dice to 0.356). These spatial rescues were achieved while demanding a median interactive workload of just 11.3% of the target volume. Acknowledging this as a simulated upper bound lacking real-world cognitive friction, the framework nevertheless demonstrates a highly Pareto-efficient pathway for safely deploying clinical AI.

Figures & tables

Explore similar work

CardsList
  1. Post-Operative Glioma Segmentation via Loss Stabilization, Normalization and Subspace Attention

    Jul 23, 2026Alexandru Crişan, Diana BorzaBrain Tumor SegmentationSwin

  2. Text-Guided Refinement of Multi-sequence Glioma Subregion Segmentation with a Vision-Language Foundation Model

    Aug 5, 2026Zach Eidex, Yu-nong Lin, Mojtaba Safari +4Brain Tumor SegmentationVision-Language Foundation Models

  3. Confidence is Not Reliability: Rethinking MC Dropout in Brain Tumour Segmentation

    Jun 17, 2026Xin Ci Wong, Duygu Sarikaya, Kieran Zucker +2Brain Tumor SegmentationMagnetic Resonance Imaging