cs.LGSep 27, 2026

A Free Knob: Decoupling Calibration and Predictive Skill in Threshold-Based Evaluation

Authors: Md Tanveer Hossain Munim, Bijoy Ahmed Saiem, Al-Amin Sany, Tanzima Hashem

Organizations: Regional Integrated Multi-Hazard Early Warning System (RIMES) · Bangladesh University of Engineering and Technology (BUET)

Abstract

Many dense-prediction benchmarks evaluate rare events by pooling prediction and target over spatial blocks, thresholding each, and scoring the contingency table. At a fixed rare operating point, the max-pooled Critical Success Index (CSI) confounds spatial discrimination with amplitude calibration: sharp observations promote many blocks above threshold, while attenuated predictions from squared-error regression leave the same blocks below it. We repurpose classical monotone calibration as a symmetric audit: a post-hoc transform fitted on held-out data and applied separately to each system. The transform cannot reverse pixel ordering, so any contrast it reproduces cannot establish improved spatial ranking. On SEVIR, two released checkpoints of one architecture differ by -29.5% in extreme-threshold CSI before the control and by +5.3% after it. Across 450 pairwise contrasts among 6 systems, the difference in pooled frequency-bias deviation is associated with how far the CSI contrast moves under the control (r = +0.796), and 51 contrasts reverse sign. At CasCast's published extreme-event operating point, the cascade-over-backbone CSI gap falls from 0.1601 to 0.0339, a 78.8% reduction; the remaining gap stays positive. The effect persists when the transform is fitted on a window before the test period, and calibration also reveals advantages hidden by a better-calibrated baseline. On geostationary infrared imagery the relative gain grows as events become rarer, crowd counting reproduces the bias-gain relationship under patch-sum pooling, and semantic segmentation, where frequency bias is already near one, shows little average change. The confound therefore requires both a fixed operating point and a training regime that leaves the output miscalibrated there. We recommend reporting pooled frequency bias and a symmetric held-out FreeKnob Audit alongside rare-event pool-and-threshold scores.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond ECE: Calibrated Size Ratio, Risk Assessment, and Confidence-Weighted Metrics

    May 3, 2026Fernando Martin-Maroto, Nabil Abderrahaman, Gonzalo G. de PolaviejaOverconfidenceConfidence Calibration

  2. CalArena: A Large-Scale Post-Hoc Calibration Benchmark

    May 28, 2026Eugène Berta, David Holzmüller, Francis Bach +1Multiclass Classification

  3. RareCP: Regime-Aware Retrieval for Efficient Conformal Prediction

    May 9, 2026Manuel Heurich, Maximilian Granz, Tim LandgrafOnline Conformal PredictionPredictive Uncertainty