Uncertainty-Guided Handshake: Efficient Human-in-the-Loop Refinement for Surgical-Grade Glioma Segmentation
Authors: Samuel Hart, Ahmad Yahya, Ahmed Karam Eldaly
Organizations: Department of Computer Science, University of Exeter, Exeter, EX4 4QF, United Kingdom. · Department of Nuclear Engineering, Faculty of Engineering, King Abdulaziz University, Jeddah, Saudi Arabia. · UCL Hawkes Institute, University College London, Gower St., London, WC1E 6AE, United Kingdom.
While state-of-the-art automated models for medical image segmentation achieve high mean performance, they frequently suffer from localized, catastrophic failures that preclude safe clinical deployment, particularly in neuro-oncology. Interactive segmentation frameworks mitigate this by incorporating human oversight, but traditionally impose prohibitive cognitive and temporal workloads by requiring clinicians to manually search for errors. In this project, we present an efficient, Hybrid Structural-Aleatoric Human-in-the-Loop framework for glioma segmentation that bridges the gap between automated baseline performance and surgical-grade precision, achieving sub-2.0 mm HD95 on curated benchmarks while providing safety-net routing for structural failures across real-world clinical data. By extracting voxel-wise Test-Time Augmentation (TTA) uncertainty and applying hierarchical topological filtering, our method proactively isolates high-risk structural anomalies. We comprehensively evaluated our approach on a challenging out-of-distribution clinical stress-test cohort (N = 362). Operating under a simulated Human Oracle, the framework improved the Whole Tumor (WT) Dice score from 0.891 to 0.914 and reduced the 95th percentile Hausdorff Distance (HD95) from 5.82 mm to 4.76 mm. Critically for surgical safety, the system rescued severe boundary failures in the Tumor Core, reducing mean HD95 from 17.96 mm to 14.83 mm (improving absolute TC Dice to 0.356). These spatial rescues were achieved while demanding a median interactive workload of just 11.3% of the target volume. Acknowledging this as a simulated upper bound lacking real-world cognitive friction, the framework nevertheless demonstrates a highly Pareto-efficient pathway for safely deploying clinical AI.
Figures & tables
Figure 1: Overview of the UG-HITL framework. Phase 1 (AI Inference): a 3D nnU-Net V2 backbone with test-time augmentation (TTA) generates the baseline segmentation and its voxel-wise predictive entropy map. Phase 2 (Decoupled Algorithmic Triage): the Track A Structural Tripwire screens each prediction for gross structural failure; collapse cases bypass refinement and are routed directly to mandatory full review, while passing cases undergo Track B refinement—topological filtering and probabilistic hotspot detection (Track B1), gradient-snapped asymmetrical search banding (Track B2), and complexity-aware dynamic budgeting capped at 25% of the target tumor volume. Phase 3 (Human-in-the-Loop): the clinician receives a prioritized queue of localized boundary proposals, and accepted proposals are merged into the final refined mask.
Methodology
Dice ↑
HD95 (mm) ↓
Workload ↓
Automated SOTA (Baseline)
0.943
3.05
100.00%
UG-HITL (MC Dropout)
0.973
1.91
0.28%
UG-HITL (5-Model Ensemble)
0.931
3.52
0.01%
UG-HITL (EDL)
0.990
0.62
1.58%
Table 1: In-Distribution Uncertainty Engine Comparison. Quantitative performance evaluated on the curated BraTS 2023 Validation Cohort. Note: Values highlighted in bold represent the optimal performance for that specific metric category, demonstrating that the proposed EDL-backed pipeline simultaneously achieves peak volumetric overlap and the lowest spatial boundary error among the active refinement methods.
Metric
Whole Tumor (WT)
Tumor Core (TC)
Enhancing Tumor (ET)
Baseline (nnU-Net)
UG-HITL (TTA)
Baseline
UG-HITL (TTA)
Baseline
UG-HITL (TTA)
Dice Score (Mean)
0.891
0.914
0.286
0.356
0.528
0.577
Dice Score (Median)
0.938
0.949
0.114
0.204
1.000
1.000
HD95 (Mean, mm)
5.82
4.76
17.96
14.83
-
-
HD95 (Median, mm)
2.82
2.23
13.52
11.42
-
-
HD95 (Max, mm)
109.15
109.15
113.34
113.34
-
-
Table 2: Quantitative evaluation of the proposed UG-HITL (TTA) framework on the out-of-distribution UTSW clinical cohort ( N=362 ). The methodology successfully constrained median human workload to 11.3% while rescuing structural boundary failures across all critical sub-regions. Bold indicates the best value per column. Dashes = undefined HD95; empty–empty Dice = 1.0; workload statistics include full-manual-review cases.
Figure 2: Reduction of 95th Percentile Hausdorff Distance (HD95) across the out-of-distribution UTSW clinical cohort (Log Scale). The framework aggressively targeted catastrophic failures in the surgical margins, reducing the mean Tumor Core (TC) HD95 from a baseline of 17.96 mm to 14.83 mm. (Note: The logarithmic scale visually compresses the upper y-axis; the absolute spatial reduction for catastrophic outliers exceeding 102 mm is highly substantial.)
Figure 3: Actual clinical workload versus the assigned workload budget for cases passing the Track A Structural Tripwire; cases routed to mandatory full review (100% workload) are excluded from this analysis. The framework scales the permitted interaction allowance up to a maximum safety threshold of 25% based on the morphological complexity (SA:V ratio) of the target tumor. Each point represents one case; the dashed line marks unity (actual workload = assigned budget), and no case exceeded its assigned allocation. The data demonstrate that the algorithm efficiently executes its spatial corrections well within these dynamically assigned limits.
Figure 4: Qualitative demonstration of Multi-Class Refinement. (A) Ground Truth annotation reveals a distinct enhancing tumor core (yellow). (B) The baseline nnU-Net entirely misses the enhancing core. (C) The UG-HITL framework successfully rescues the internal core boundaries, drastically reducing the surgical margin error.
Figure 5: Qualitative demonstration of an algorithmic "Catastrophic Rescue". (Left to Right): 1. Raw MRI. 2. Baseline hallucination (red). 3. Isolated uncertainty (yellow). 4. Refined Handshake (green). Note: This sample illustrates the foundational geometric logic used by the finalised TTA framework to achieve the spatial improvements reported on the UTSW cohort.
Pipeline Configuration
Dice (WT) ↑
HD95 (WT) ↓
Dice (TC) ↑
HD95 (TC) ↓
Workload (%) ↓
1. Baseline TTA (No Refinement)
0.884
7.08 mm
0.315
20.59 mm
11.68%
2. + Hotspot Hunter Only
0.884
7.07 mm
0.315
20.46 mm
11.87%
3. + Topological & Anomaly Filtering
0.896
6.71 mm
0.369
17.66 mm
18.92%
4. + Gradient Snapping (Full Pipeline)
0.896
6.85 mm
0.368
17.76 mm
18.89%
Table 3: Ablation study isolating the quantitative impact of the framework’s individual triage components. Evaluated on a 50-case subset of the UTSW clinical cohort. (Note: Baseline metrics differ slightly from the full cohort in Table 2 due to evaluation isolated to this specific representative subset).
Department of Biomedical Informatics, Emory University School of Medicine, Atlanta, GA, USA · Department of Radiation and Cellular Oncology, The University of Chicago, Chicago, IL, USA · Department of Radiation Oncology, Winship Cancer Institute, Emory University School of Medicine, Atlanta, GA, USA
Centre for Doctoral Training in AI for Medical Diagnosis and Care, School of Computing, University of Leeds · School of Computer Science, University of Leeds · Leeds Cancer Centre, St James’s University Hospital, Leeds, UK