Many dense-prediction benchmarks evaluate rare events by pooling prediction and target over spatial blocks, thresholding each, and scoring the contingency table. At a fixed rare operating point, the max-pooled Critical Success Index (CSI) confounds spatial discrimination with amplitude calibration: sharp observations promote many blocks above threshold, while attenuated predictions from squared-error regression leave the same blocks below it. We repurpose classical monotone calibration as a symmetric audit: a post-hoc transform fitted on held-out data and applied separately to each system. The transform cannot reverse pixel ordering, so any contrast it reproduces cannot establish improved spatial ranking. On SEVIR, two released checkpoints of one architecture differ by -29.5% in extreme-threshold CSI before the control and by +5.3% after it. Across 450 pairwise contrasts among 6 systems, the difference in pooled frequency-bias deviation is associated with how far the CSI contrast moves under the control (r = +0.796), and 51 contrasts reverse sign. At CasCast's published extreme-event operating point, the cascade-over-backbone CSI gap falls from 0.1601 to 0.0339, a 78.8% reduction; the remaining gap stays positive. The effect persists when the transform is fitted on a window before the test period, and calibration also reveals advantages hidden by a better-calibrated baseline. On geostationary infrared imagery the relative gain grows as events become rarer, crowd counting reproduces the bias-gain relationship under patch-sum pooling, and semantic segmentation, where frequency bias is already near one, shows little average change. The confound therefore requires both a fixed operating point and a training regime that leaves the output miscalibrated there. We recommend reporting pooled frequency bias and a symmetric held-out FreeKnob Audit alongside rare-event pool-and-threshold scores.
Figures & tables
Figure 1: One SEVIR event. The forecast broadly follows the observed storm but attenuates its peaks. After 16×16 max pooling and thresholding at 219 , the observation has 94 positive blocks and the forecast 16, illustrating the large deficit in forecast exceedances. Its block ratio, 16/94≈0.17 , is the closest of any frame to the backbone’s pooled frequency bias over the whole split (0.172), so the event shows the backbone’s typical deficit.
comparison
bias ref
bias new
raw
calibrated
rel.-gain reduction
cascade vs. backbone
0.172
1.037
+109.1%
+12.9%
88%
cascade vs. EarthFormer
0.281
1.037
+47.4%
+18.9%
60%
EarthFormer vs. persistence
1.020
0.281
+8.2%
+31.3%
↑
Table 1: Our evaluation of the three model pairs compared in prior work, at max 16×16 pooling and τ=219 . “Raw” is the relative CSI gain reported without calibration; “calibrated” is the same relative gain after the global held-out quantile map is fitted separately to both arms. The final column is the reduction in that relative gain; ↑ marks a gain that grows. Bias columns give the uncalibrated pooled frequency bias of each arm.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
metric
pooling
published
ours
deviation
CSI-M
none
0.4310
0.4299
-0.2%
CSI-M
max 162
0.4351
0.4327
-0.5%
CSI-219
none
0.1448
0.1488
+2.8%
CSI-219
max 162
0.1481
0.1467
-0.9%
Appendix
Table 2: Reproduction of the CasCast deterministic backbone against the values published in the authors’ Table 4. CSI-M is the mean over the six thresholds.
class
R
Bias
IoU
Δ
class
R
Bias
IoU
Δ
person
9.012%
1.022
0.8155
-0.0008
tvmonitor
0.414%
0.963
0.6992
-0.0248
diningtable
3.002%
1.732
0.3802
-0.0351
horse
0.336%
0.997
0.6969
+0.0010
bus
0.868%
1.005
0.7802
-0.0123
cow
0.299%
0.964
0.7718
+0.0000
train
0.788%
0.975
0.7764
+0.0000
sheep
0.258%
0.988
0.7033
-0.0143
chair
0.729%
0.643
0.3384
+0.0204
aeroplane
0.227%
0.999
0.8153
+0.0000
sofa
0.702%
0.825
0.5066
+0.0226
bottle
0.216%
0.780
0.5125
+0.0207
Appendix
Table 3: COCO val2017 , deeplabv3-resnet50, released checkpoint. Each class is evaluated independently as a binary mask at probability threshold 0.5 ; per-class IoU ≡CSI is reported at that operating point and after one monotone logit shift δ fitted on the disjoint calibration half by ∣Bias−1∣ . This audit is distinct from the separate multiclass argmax reproduction check. Classes are ordered by positive-pixel fraction; those without enough calibration support to fit a knob are omitted rather than fitted on noise.
τ
R (test)
CSI identity
CSI recal
ΔCSI
rel. gain
n
ΔCSI (all)
median (all)
220
20.270%
0.2672
0.2610
-0.0062
-2.3%
10/10
-0.0062
-0.0048
215
11.537%
0.1992
0.1956
-0.0036
-1.8%
10/10
-0.0036
-0.0011
210
6.378%
0.1381
0.1368
-0.0013
-0.9%
10/10
-0.0013
-0.0039
205
3.124%
0.0578
0.0739
+0.0160
+27.7%
10/10
+0.0160
+0.0143
200
0.884%
0.0117
0.0352
+0.0236
+202.1%
9/10
+0.0238
+0.0245
197
0.264%
0.0085
0.0299
+0.0214
+250.8%
7/10
+0.0237
+0.0265
Appendix
Table 4: Gain from a free monotone recalibration against event rarity. R is the all-scene test fraction on the valid evaluation footprint; scores retain the observed-coverage >0.1% frame filter. The relative gain is undefined for a run whose uncalibrated CSI is zero, so n gives the runs entering that column against the 10 trained at that τ . Columns three to six are means over those n runs, which is why ΔCSI and ΔCSI (all) differ at the two rarest levels; the last two columns are computed over all 10. The 4 runs the relative column omits score exactly zero uncalibrated and are lifted to between 0.0258 and 0.0315 CSI by the knob, so their true relative gain is unbounded and dropping them can only understate the trend.
pooling
Bias
CSI
CSI+ qmap
gain
pysteps > EarthFormer
uncal.
+global
+per-cell
none
0.291
0.1488
0.2074
+39.3%
0
0
0
avg 42
0.350
0.1724
0.2185
+26.7%
0
0
0
max 42
0.214
0.1426
0.2297
+61.1%
0
0
0
avg 162
0.664
0.2347
0.2138
-8.9%
0
0
0
max 162
0.172
0.1467
0.2639
+79.8%
5
5
0
Appendix
Table 5: Pooling convention determines the size of the calibration effect and the ranking. Bias and CSI are for the CasCast deterministic backbone at τ=219 . The last three columns count, out of six thresholds, how often pysteps (2019 advection) outranks EarthFormer (2022) when both are left uncalibrated, both receive the global quantile map, and both receive the per-cell knob.
arm
pooled Bias
ΔCSI from the knob
τ=3
τ=7
τ=10
τ=3
τ=7
τ=10
CSRNet
1.122
0.682
0.289
+0.0072
+0.0352
+0.0911
CNN baseline (MSE)
1.000
1.128
1.426
+0.0021
-0.0339
-0.0442
target, displaced 2 blocks
0.938
0.954
0.969
-0.0009
+0.0000
+0.0000
target, blurred (mass conserved)
0.995
0.975
0.976
+0.0000
+0.0000
+0.0000
target, peaks compressed (mass conserved)
0.942
0.831
0.761
+0.0047
+0.0000
+0.0809
Appendix
Table 6: Crowd density under the patch-sum protocol of Section 7 ( 32×32 patch sums), averaged over 5 seeds. ΔCSI is the absolute change from one monotone knob, selected on a held-out split by ∣Bias−1∣ and never on CSI . The last two rows are targets we corrupted deliberately, to move placement and sharpness one at a time. Selecting on ∣Bias−1∣ cannot tune for the score we report, but it is not neutral toward an arm whose bias is long; re-selecting the same family on validation CSI removes the harm and leaves every conclusion unchanged (Appendix J ).
seed
pooled Bias
CSI
patch MAE
bg MAE
s42
0.002
0.0017
8.12
0.794
s43
0.107
0.0777
6.23
0.806
s44
0.136
0.0687
6.38
0.885
s45
1.146
0.3194
4.72
0.798
s46
0.053
0.0368
6.61
1.017
Appendix
Table 7: CSRNet on ShanghaiTech Part A, 5 seeds, one recipe, at τ=10 under the patch-sum protocol. “patch MAE” is the mean absolute count error over patches above the threshold and “bg MAE” over the rest. Only the seed changes across rows.
τ
bias EF
bias CasCast
raw ΔCSI
cal. ΔCSI
raw
calibrated
16
0.915
0.875
-0.0109
+0.0065
-1.4%
+0.9%
74
0.747
0.718
-0.0149
-0.0007
-2.2%
-0.1%
133
0.534
0.495
-0.0236
+0.0055
-5.1%
+1.1%
160
0.387
0.337
-0.0315
+0.0054
-9.5%
+1.4%
181
0.337
0.285
-0.0324
+0.0118
-11.3%
+3.4%
219
0.281
0.172
-0.0614
+0.0133
-29.5%
+5.3%
Appendix
Table 8: One architecture, two independent trainings, max 16×16 pooling. “raw” columns are what a results table would print; “calibrated” columns follow one global quantile map fitted per arm on the calibration half. The sign of the contrast reverses at 5 of the 6 thresholds.
arm
Bias
CSI
CSI test-fit
CSI val-fit
gain test-fit
gain val-fit
persistence
1.020
0.1924
0.1908
0.1936
-0.8%
+0.6%
pysteps
0.907
0.2700
0.2704
0.2719
+0.2%
+0.7%
EarthFormer
0.281
0.2082
0.2505
0.2877
+20.3%
+38.2%
CasCast backbone
0.172
0.1467
0.2639
0.2928
+79.8%
+99.6%
CasCast cascade
1.037
0.3069
0.2978
0.3063
-2.9%
-0.2%
CasCast cascade (no guidance)
0.548
0.2475
0.2899
0.2662
+17.1%
+7.6%
Appendix
Table 9: The same arms at max 16×16 pooling and τ=219 , scored on the test evaluation half, with the global quantile map fitted two ways. “test-fit” is the paper’s default, on the disjoint calibration half of the test split; “val-fit” is on 128 events from 2019-01-01 to 2019-04-30, which precede the test period and were held out of both models’ training. The backbone’s val-fit score comes from the earlier inference run described in Table 12 .
CSI
ETS
τ
raw
calibrated
rel.-gain reduction
raw
calibrated
rel.-gain reduction
16
-0.2%
+4.9%
—
-1.8%
+5.0%
—
74
+7.4%
+4.7%
36%
+7.4%
+4.4%
41%
133
+26.9%
+12.2%
55%
+27.0%
+11.9%
56%
160
+52.4%
+14.8%
72%
+52.0%
+14.3%
73%
181
+64.7%
+14.5%
78%
+64.1%
+14.0%
78%
Appendix
Table 10: The cascade-over-backbone gain of Section 6 at max 16×16 , raw and after both arms receive the same monotone knob, scored under CSI and under ETS. “rel.-gain reduction” is the fraction of the raw relative gain removed by the knob.
comparison
τ
bias ref
bias new
raw
calibrated
rel.-gain reduction
cascade vs. own backbone
74
0.718
0.961
+7.4%
+4.7%
36%
133
0.495
0.905
+26.9%
+12.2%
55%
160
0.337
0.875
+52.4%
+14.8%
72%
181
0.285
0.904
+64.7%
+14.5%
78%
219
0.172
1.037
+109.1%
+12.9%
88%
cascade vs. EarthFormer
133
0.534
0.905
+20.5%
+13.5%
34%
Appendix
Table 11: Full threshold-by-threshold version of the three contrasts summarised in the main text, at max 16×16 pooling. The first block runs from τ=74 because the relative-gain reduction quoted in the text starts there; the others start at 133 . “raw” is the relative gain as a results table would print it, “calibrated” the same gain after one global quantile map is fitted per arm on the calibration half and applied to both, and “rel.-gain reduction” is 1−calibrated/raw , with ↑ marking rows where calibration raises the gain so that ratio has no stable base. In each label “X vs. Y”, Y is the reference arm and X the one claimed better, and the bias columns are their uncalibrated pooled biases in that order.
τ
CSI (point estimate)
knob gain
residual
backbone
+cal
cascade
+cal
backbone+cal − backbone
cascade+cal − backbone+cal
16
0.7850
0.7471
0.7833
0.7833
-0.0379 [-0.041, -0.035]
+0.0363 [+0.031, +0.041]
74
0.6654
0.6829
0.7146
0.7151
+0.0175 [+0.016, +0.018]
+0.0323 [+0.028, +0.036]
133
0.4442
0.4968
0.5638
0.5574
+0.0525 [+0.049, +0.055]
+0.0606 [+0.055, +0.066]
160
0.3013
0.3962
0.4592
0.4550
+0.0949 [+0.090, +0.100]
+0.0587 [+0.053, +0.065]
181
0.2536
0.3612
0.4176
0.4134
+0.1076 [+0.103, +0.112]
+0.0522 [+0.046, +0.058]
Appendix
Table 12: Paired bootstrap over the 936 evaluation events, 2000 replicates, max 16×16 pooling. Brackets are 95% intervals on the paired contrast; the four CSI columns are point estimates. Column six is what the free knob is worth to the backbone; column seven is what remains of the cascade’s advantage once both arms carry that same transform. Both contrasts have p<0.001 at every threshold. Point estimates are those of Section 6 . The intervals resample per-event counts cached from an earlier inference run of the deterministic backbone, whose CSI differs from the current run by at most 1.1×10−4 in any cell; the counts of every other arm are identical in the two runs.
pooled bias
CSI
lead
id
G
L
id
G
L
1
0.572
1.158
0.703
0.5000
0.5952
0.5605
2
0.397
0.882
0.539
0.3466
0.5208
0.4344
3
0.262
0.652
0.435
0.2282
0.4167
0.3324
4
0.187
0.508
0.406
0.1604
0.3279
0.2894
5
0.144
0.382
0.367
0.1209
0.2559
0.2485
Appendix
Table 13: Pooled frequency bias and CSI by lead time at max 16×16 , τ=219 , for CasCast’s deterministic backbone. “id” is uncalibrated, “G” the single global quantile map, “L” the map fitted per lead time at the equal-pixel budget. The global map overshoots at the first lead and under-corrects at the last; the per-lead map is flatter in bias and scores higher at long leads, but the tables are summed across leads and the early leads carry the hits.
family
rule
backbone
cascade
residual
rel.-gain reduction
pooled-horizon
bias
0.3201
0.3059
-4.4%
104.0%
pooled-horizon
val-CSI
0.3201
0.3123
-2.4%
102.2%
per-horizon
bias
0.3366
0.2990
-11.2%
110.0%
per-horizon
val-CSI
0.3413
0.3086
-9.6%
108.6%
Appendix
Table 14: The knob family crossed against the selection rule, at max 16×16 , τ=219 , decomposing the cascade contrast whose raw gain is +111.3%. Both rules see the calibration half only. “per-horizon” selects a member per (cell, lead) and is not a member of G ; it is reported as an additional sensitivity result. The range in Section 6 compares its two specified controls only.