Beyond scalar losses: calibrating segmentation models via gradient vector field surgery
Authors: Laurin Lux, Alexander H. Berger, Moritz Knolle, Daniel Rückert, Johannes C. Paetzold
Organizations: School of Computation, Information and Technology, TUM, Munich, Germany · Munich Center for Machine Learning, Munich, Germany · Department of Radiology, Weill Cornell Medicine, New York City, USA · School of Medicine and Health, TUM University Hospital, Munich, Germany · Department of Computing, Imperial College London, London, UK · Cornell Tech, New York City, USA
Region-based loss functions, such as the Dice loss, have established themselves as the de facto standard for highly class- and region-imbalanced segmentation tasks. However, models trained using region-based loss functions are notoriously miscalibrated and typically yield over-confident predictions. In medical imaging applications, such as defining tumor resection margins, this miscalibration is hindering clinical adoption. In this work, we outline a novel gradient perspective on this overconfidence and show how it affects region-based loss functions. We propose a "surgery" on the gradient vector field as a simple, yet effective intervention to mitigate calibration issues. This surgery adds a factor to the loss's partial derivative, scaling the gradient's magnitude linearly with the prediction error. In empirical evaluations across 2D and 3D medical segmentation tasks, we demonstrate the effectiveness of this intervention while maintaining high prediction accuracy when used in conjunction with any region-based loss function.
The Segment Anything Model (SAM) exhibits strong zero-shot performance on natural images but suffers from domain shift and overconfidence when applied to medical volumes. We propose \textbf{CalSAM}, a lightweight adaptation framework that (i) reduces encoder sensitivity to domain shift via a \emph{Feature Fisher Information Penalty} (FIP) computed on 3D feature maps and (ii) penalizes overconfident voxel-wise errors through a \emph{Confidence Misalignment Penalty} (CMP). The combined loss, LCalSAM fine-tunes only the mask decoder while keeping SAM's encoders frozen. On cross-center and scanner-shift evaluations, CalSAM substantially improves accuracy and calibration: e.g., on the BraTS scanner split (Siemens→GE) CalSAM shows a +7.4% relative improvement in DSC (80.1% vs.\ 74.6%), a −26.9% reduction in HD95 (4.6 mm vs.\ 6.3 mm), and a −39.5% reduction in ECE (5.2% vs.\ 8.6%). On ATLAS-C (motion corruptions), CalSAM achieves a +5.3% relative improvement in DSC (75.9%) and a −32.6% reduction in ECE (5.8%). Ablations show FIP and CMP contribute complementary gains (p<0.01), and the Fisher penalty incurs a modest ∼15% training-time overhead. CalSAM therefore delivers improved domain generalization and better-calibrated uncertainty estimates for brain MRI segmentation, while retaining the computational benefits of freezing SAM's encoder.
Behraj Khan, Tahir Qasim Syed, Syed Ahmad Chan Bukhari
Medical experts often manually segment images to obtain diagnostic statistics and discard the resulting annotations. We aim to train segmentation models to alleviate this burden, but constrained to the retained summary statistics (e.g., the area of the annotated region). Empirical results suggest that statistics alone are insufficient for this task, but adding weak information in the form of a few pixels within the area of interest significantly improves performance. We use a novel loss function that combines terms for image reconstruction quality, matching to summary statistics, and overlap between the predicted foreground and the weak supervisory signal. Experiments on standard image, ultrasound (breast cancer), and Computed Tomography (CT) scan (kidney tumors) data demonstrate the utility and potential of the approach.
Reliable confidence estimates are essential in semantic segmentation, yet modern models often remain miscalibrated. We investigate two overlooked issues in post-hoc calibration. First, adding a constant to all logits leaves softmax probabilities unchanged, but several standard calibrators depend on this arbitrary offset. In segmentation, this offset can vary across pixels or voxels, introducing spatially varying representation dependence. We characterize translation-invariant (TI) calibrators and construct TI counterparts of shift-sensitive methods. Second, calibrating with cross-entropy can degrade segmentation quality due to mismatched training and calibration objectives and limited calibration data. We investigate decision-preserving calibration under argmax- and order-preservation constraints. Since these constraints restrict affine softmax calibrators to temperature scaling, we introduce more expressive class-conditional affine calibrators that preserve decisions. Across natural-image and medical segmentation benchmarks, including corruption-based covariate shift, TI variants generally improve calibration, while decision-preserving variants prevent segmentation degradation by construction and retain strong calibration performance. Our findings provide practical design principles for post-hoc calibration in semantic segmentation.