StatD2GAN: When Calibration Masks Generator Quality in Held-Out Evaluation of Synthetic Weather Sequences
Authors: Mustafa Ozaytac, Ozge Karadag Atas
Organizations: Department of Statistics, Graduate School of Science and Engineering, Hacettepe University, Ankara, 06800, Turkey · Department of Statistics, Faculty of Science, Hacettepe University, Ankara, 06800, Turkey
Generative models for multivariate weather series are routinely evaluated with pooled distributional metrics computed after marginal calibration. We show this practice can invalidate architectural conclusions, and rebuild the evaluation of StatD2GAN, a three-discriminator GAN with evolutionary weight adaptation, around a held-out protocol: the final two calendar years of each dataset are held out behind a 168 hour embargo, calibration is fitted on the training block only, and all metrics are computed on the held-out block. Evidence comes from 25 matched (location, seed) pairs across five Koppen-Geiger climates, tested with Wilcoxon signed-rank tests under Holm correction. Four results follow. First, isotonic calibration drives the Kolmogorov-Smirnov distance to within 2% of a per-location noise-and-shift floor for every architecture tested, including a deliberately weak RCGAN baseline, so calibrated marginal metrics cannot discriminate between architectures. Second, the sorted-representation discriminator is the only component whose removal significantly degrades cross-variable dependence (Kendall tau MAE +0.080, Holm p = 0.009), with a regime-dependent effect: near zero in Ankara, above 115% in Dubai and Yakutsk. A rank-transformed variant isolates the mechanism as quantile supervision of the marginals rather than copula matching. Third, physical constraint violations are injected by calibration, not the generator; projection removes them at negligible cost (deltaKS <= 0.003). Fourth, pooled metrics conceal a collapse of between-sequence weekly-mean variability, a proxy for seasonal and regime diversity, in TimeGAN that only sequence-level statistics expose. We recommend floor-referenced marginal evaluation, matched-pair testing, and sequence-level variance decomposition as minimum requirements for calibrated generative pipelines.
Figures & tables
Location
Köppen
Climate regime
Primary challenge
Ankara, Turkey
BSk
Semi-arid continental
Low humidity, moderate variability
Dubai, UAE
BWh
Hot desert
Extreme aridity, humidity saturation
Bergen, Norway
Cfb
Oceanic
High precipitation, complex coupling
Yakutsk, Russia
Dfd
Subarctic
Extreme cold, widest seasonal range
Lhasa, China
Dwb
High-altitude monsoonal
Altitude effects; added post hoc
Table 1: ERA5 datasets and climate classifications.
Category
Parameter
Value
Data
Sequence length T
168 (hourly, stride 24)
Features F
11 (7 scaled to [−1,1] )
Split
last 2 calendar years held out, 168 h embargo
Noise dimension dz
128
Model
Hidden dimension
512
λ(0)
[1,1,1]
Table 2: Model and training hyperparameters.
Location
Ankara
Bergen
Dubai
Lhasa
Yakutsk
Ablation arm
Full model
0.161 ± 0.058
0.206 ± 0.089
0.159 ± 0.085
0.202 ± 0.071
0.135 ± 0.015
No Dtemp
0.161 ± 0.063
0.211 ± 0.075
0.14 ± 0.042
0.178 ± 0.019
0.096 ± 0.028
No Dstat
0.226 ± 0.077
0.18 ± 0.036
0.138 ± 0.041
0.176 ± 0.047
0.124 ± 0.048
No Dsort
0.161 ± 0.077
0.253 ± 0.039
0.383 ± 0.156
0.251 ± 0.057
0.29 ± 0.136
No calibration
0.213 ± 0.082
0.227 ± 0.089
0.171 ± 0.024
0.142 ± 0.024
0.132 ± 0.046
Table 3: Kendall tau MAE (mean ± SD, N=5 seeds, held-out).
Location
Ankara
Bergen
Dubai
Lhasa
Yakutsk
Ablation arm
Full model
0.332 ± 0.057
0.34 ± 0.042
0.28 ± 0.081
0.253 ± 0.056
0.31 ± 0.045
No Dtemp
0.307 ± 0.044
0.328 ± 0.094
0.291 ± 0.039
0.314 ± 0.023
0.365 ± 0.015
No Dstat
0.333 ± 0.107
0.357 ± 0.033
0.326 ± 0.062
0.304 ± 0.087
0.309 ± 0.067
No Dsort
0.36 ± 0.094
0.387 ± 0.075
0.251 ± 0.061
0.243 ± 0.02
0.319 ± 0.117
No calibration
0.325 ± 0.032
0.364 ± 0.073
0.279 ± 0.061
0.206 ± 0.024
0.34 ± 0.112
Table 4: ACF MAE (mean ± SD, N=5 seeds, held-out).
Table 7: Real train–val KS noise floor per location.
Metric
Contrast (vs. full model)
Median Δ
pWilcoxon
Sign count (+)
pHolm
KS sup-distance
No calibration
0.1105
<0.0001
AN:5/5 DU:5/5 BE:5/5 LH:5/5 YA:5/5
<0.0001
KS sup-distance
No Dtemp
-0.0001
0.0081
AN:1/5 DU:1/5 BE:2/5 LH:0/5 YA:1/5
0.0565
KS sup-distance
Moment Dstat
-0.0001
0.0551
AN:2/5 DU:1/5 BE:2/5 LH:1/5 YA:2/5
0.3304
KS sup-distance
Rank Dsort
0.0001
0.0626
AN:3/5 DU:3/5 BE:4/5 LH:4/5 YA:2/5
0.3131
KS sup-distance
No Dstat
0.0000
0.2002
AN:2/5 DU:1/5 BE:3/5 LH:1/5 YA:2/5
0.8009
KS sup-distance
No auxiliary losses
-0.0001
0.2002
AN:2/5 DU:1/5 BE:3/5 LH:1/5 YA:2/5
0.6006
Table 8: Paired Wilcoxon over 25 (location, seed) pairs; Holm within metric. Sign count (+): number of seeds, out of 5, for which the ablation arm’s value exceeded the full model’s, per location (AN = Ankara, DU = Dubai, BE = Bergen, LH = Lhasa, YA = Yakutsk). Contrasts are sorted by pWilcoxon within each metric; ablation-arm definitions are given in Table 3 .
Location
PhysViol before
PhysViol after
Δ KS
Δτ -MAE
Ankara
0.045
0.0
+0.0010
−0.0033
Dubai
0.052
0.0
+0.0029
−0.0069
Bergen
0.044
0.0
+0.0003
−0.0001
Lhasa
0.060
0.0
+0.0015
−0.0038
Yakutsk
0.038
0.0
+0.0000
+0.0006
Table 9: Constraint projection on calibrated full-model output (mean over 5 seeds per location; source: projection_v2 ). Violations drop to exactly zero in every run; the distributional cost is negligible against run-to-run noise ( ∣Δτ∣≈0.085 ).
KS (raw)
KS (calib.)
Kendall τ MAE
ACF MAE
Phys. viol. (raw)
Phys. viol. rate
Mean
SD
Mean
SD
Mean
SD
Mean
SD
Mean
SD
Mean
SD
Location
Model
Ankara
RCGAN
0.752
0.090
0.054
0.039
0.445
0.142
0.265
0.054
0.196
0.433
0.250
0.039
TimeGAN
0.143
0.030
0.036
0.000
0.238
0.186
0.186
0.065
0.001
0.003
0.014
0.011
Bergen
RCGAN
0.713
0.048
0.036
0.000
0.405
0.199
0.245
0.059
0.514
0.498
0.322
0.100
TimeGAN
0.141
0.020
0.036
0.000
0.145
0.033
0.239
0.081
0.044
0.055
0.110
0.051
Table 10: Baselines under the held-out protocol (mean ± SD, 5 seeds). Raw columns are pre-calibration; calibrated KS locks to the per-location noise floor for all models.
Location
Model
Between-seq. SD (K)
Within-seq. SD (K)
Step diff (K)
Ankara
Real (val)
8.21
4.44
0.94
StatD2GAN
6.25
4.35
1.30
RCGAN
0.16
2.31
0.21
TimeGAN
0.03
9.81
1.78
Bergen
Real (val)
5.18
2.14
0.32
StatD2GAN
4.12
2.01
0.53
Table 11: Sequence-level temperature statistics (mean over 5 seeds). Between-seq. SD is the standard deviation of per-sequence means, i.e. seasonal diversity; within-seq. SD is the mean within-sequence standard deviation. Pooled metrics (Tables 3 – 6 ) are blind to both. Real (val): held-out real data; RCGAN and TimeGAN are the baselines of Table 10 .