StatD2GAN: When Calibration Masks Generator Quality in Held-Out Evaluation of Synthetic Weather Sequences
Authors: Mustafa Ozaytac, Ozge Karadag Atas
Organizations: Department of Statistics, Graduate School of Science and Engineering, Hacettepe University, Ankara, 06800, Turkey · Department of Statistics, Faculty of Science, Hacettepe University, Ankara, 06800, Turkey
Generative models for multivariate weather series are routinely evaluated with pooled distributional metrics computed after marginal calibration. We show this practice can invalidate architectural conclusions, and rebuild the evaluation of StatD2GAN, a three-discriminator GAN with evolutionary weight adaptation, around a held-out protocol: the final two calendar years of each dataset are held out behind a 168 hour embargo, calibration is fitted on the training block only, and all metrics are computed on the held-out block. Evidence comes from 25 matched (location, seed) pairs across five Koppen-Geiger climates, tested with Wilcoxon signed-rank tests under Holm correction. Four results follow. First, isotonic calibration drives the Kolmogorov-Smirnov distance to within 2% of a per-location noise-and-shift floor for every architecture tested, including a deliberately weak RCGAN baseline, so calibrated marginal metrics cannot discriminate between architectures. Second, the sorted-representation discriminator is the only component whose removal significantly degrades cross-variable dependence (Kendall tau MAE +0.080, Holm p = 0.009), with a regime-dependent effect: near zero in Ankara, above 115% in Dubai and Yakutsk. A rank-transformed variant isolates the mechanism as quantile supervision of the marginals rather than copula matching. Third, physical constraint violations are injected by calibration, not the generator; projection removes them at negligible cost (deltaKS <= 0.003). Fourth, pooled metrics conceal a collapse of between-sequence weekly-mean variability, a proxy for seasonal and regime diversity, in TimeGAN that only sequence-level statistics expose. We recommend floor-referenced marginal evaluation, matched-pair testing, and sequence-level variance decomposition as minimum requirements for calibrated generative pipelines.
Figures & tables
Location
Köppen
Climate regime
Primary challenge
Ankara, Turkey
BSk
Semi-arid continental
Low humidity, moderate variability
Dubai, UAE
BWh
Hot desert
Extreme aridity, humidity saturation
Bergen, Norway
Cfb
Oceanic
High precipitation, complex coupling
Yakutsk, Russia
Dfd
Subarctic
Extreme cold, widest seasonal range
Lhasa, China
Dwb
High-altitude monsoonal
Altitude effects; added post hoc
Table 1: ERA5 datasets and climate classifications.
Category
Parameter
Value
Data
Sequence length T
168 (hourly, stride 24)
Features F
11 (7 scaled to [−1,1] )
Split
last 2 calendar years held out, 168 h embargo
Noise dimension dz
128
Model
Hidden dimension
512
λ(0)
[1,1,1]
Table 2: Model and training hyperparameters.
Location
Ankara
Bergen
Dubai
Lhasa
Yakutsk
Ablation arm
Full model
0.161 ± 0.058
0.206 ± 0.089
0.159 ± 0.085
0.202 ± 0.071
0.135 ± 0.015
No Dtemp
0.161 ± 0.063
0.211 ± 0.075
0.14 ± 0.042
0.178 ± 0.019
0.096 ± 0.028
No Dstat
0.226 ± 0.077
0.18 ± 0.036
0.138 ± 0.041
0.176 ± 0.047
0.124 ± 0.048
No Dsort
0.161 ± 0.077
0.253 ± 0.039
0.383 ± 0.156
0.251 ± 0.057
0.29 ± 0.136
No calibration
0.213 ± 0.082
0.227 ± 0.089
0.171 ± 0.024
0.142 ± 0.024
0.132 ± 0.046
Table 3: Kendall tau MAE (mean ± SD, N=5 seeds, held-out).
Location
Ankara
Bergen
Dubai
Lhasa
Yakutsk
Ablation arm
Full model
0.332 ± 0.057
0.34 ± 0.042
0.28 ± 0.081
0.253 ± 0.056
0.31 ± 0.045
No Dtemp
0.307 ± 0.044
0.328 ± 0.094
0.291 ± 0.039
0.314 ± 0.023
0.365 ± 0.015
No Dstat
0.333 ± 0.107
0.357 ± 0.033
0.326 ± 0.062
0.304 ± 0.087
0.309 ± 0.067
No Dsort
0.36 ± 0.094
0.387 ± 0.075
0.251 ± 0.061
0.243 ± 0.02
0.319 ± 0.117
No calibration
0.325 ± 0.032
0.364 ± 0.073
0.279 ± 0.061
0.206 ± 0.024
0.34 ± 0.112
Table 4: ACF MAE (mean ± SD, N=5 seeds, held-out).
Table 7: Real train–val KS noise floor per location.
Metric
Contrast (vs. full model)
Median Δ
pWilcoxon
Sign count (+)
pHolm
KS sup-distance
No calibration
0.1105
<0.0001
AN:5/5 DU:5/5 BE:5/5 LH:5/5 YA:5/5
<0.0001
KS sup-distance
No Dtemp
-0.0001
0.0081
AN:1/5 DU:1/5 BE:2/5 LH:0/5 YA:1/5
0.0565
KS sup-distance
Moment Dstat
-0.0001
0.0551
AN:2/5 DU:1/5 BE:2/5 LH:1/5 YA:2/5
0.3304
KS sup-distance
Rank Dsort
0.0001
0.0626
AN:3/5 DU:3/5 BE:4/5 LH:4/5 YA:2/5
0.3131
KS sup-distance
No Dstat
0.0000
0.2002
AN:2/5 DU:1/5 BE:3/5 LH:1/5 YA:2/5
0.8009
KS sup-distance
No auxiliary losses
-0.0001
0.2002
AN:2/5 DU:1/5 BE:3/5 LH:1/5 YA:2/5
0.6006
Table 8: Paired Wilcoxon over 25 (location, seed) pairs; Holm within metric. Sign count (+): number of seeds, out of 5, for which the ablation arm’s value exceeded the full model’s, per location (AN = Ankara, DU = Dubai, BE = Bergen, LH = Lhasa, YA = Yakutsk). Contrasts are sorted by pWilcoxon within each metric; ablation-arm definitions are given in Table 3 .
Location
PhysViol before
PhysViol after
Δ KS
Δτ -MAE
Ankara
0.045
0.0
+0.0010
−0.0033
Dubai
0.052
0.0
+0.0029
−0.0069
Bergen
0.044
0.0
+0.0003
−0.0001
Lhasa
0.060
0.0
+0.0015
−0.0038
Yakutsk
0.038
0.0
+0.0000
+0.0006
Table 9: Constraint projection on calibrated full-model output (mean over 5 seeds per location; source: projection_v2 ). Violations drop to exactly zero in every run; the distributional cost is negligible against run-to-run noise ( ∣Δτ∣≈0.085 ).
KS (raw)
KS (calib.)
Kendall τ MAE
ACF MAE
Phys. viol. (raw)
Phys. viol. rate
Mean
SD
Mean
SD
Mean
SD
Mean
SD
Mean
SD
Mean
SD
Location
Model
Ankara
RCGAN
0.752
0.090
0.054
0.039
0.445
0.142
0.265
0.054
0.196
0.433
0.250
0.039
TimeGAN
0.143
0.030
0.036
0.000
0.238
0.186
0.186
0.065
0.001
0.003
0.014
0.011
Bergen
RCGAN
0.713
0.048
0.036
0.000
0.405
0.199
0.245
0.059
0.514
0.498
0.322
0.100
TimeGAN
0.141
0.020
0.036
0.000
0.145
0.033
0.239
0.081
0.044
0.055
0.110
0.051
Table 10: Baselines under the held-out protocol (mean ± SD, 5 seeds). Raw columns are pre-calibration; calibrated KS locks to the per-location noise floor for all models.
Location
Model
Between-seq. SD (K)
Within-seq. SD (K)
Step diff (K)
Ankara
Real (val)
8.21
4.44
0.94
StatD2GAN
6.25
4.35
1.30
RCGAN
0.16
2.31
0.21
TimeGAN
0.03
9.81
1.78
Bergen
Real (val)
5.18
2.14
0.32
StatD2GAN
4.12
2.01
0.53
Table 11: Sequence-level temperature statistics (mean over 5 seeds). Between-seq. SD is the standard deviation of per-sequence means, i.e. seasonal diversity; within-seq. SD is the mean within-sequence standard deviation. Pooled metrics (Tables 3 – 6 ) are blind to both. Real (val): held-out real data; RCGAN and TimeGAN are the baselines of Table 10 .
Training robust multivariate time series forecasting models requires large, diverse corpora, yet many real-world domains provide only a handful of observed sequences. Existing generators fail to resolve this mismatch: prior-based approaches (e.g., CauKer, TimePFN) produce domain-agnostic samples, while data-driven methods (e.g., TimeGAN) treat references as black-box supervision, forfeiting explicit control over periodic structure, local variability, and cross-variable dynamics. We propose ReGeN, a reference-guided generative pipeline that treats observed sequences not as examples to imitate, but as structural scaffolds for controllable synthesis. ReGeN decomposes each reference into three interpretable components: a phase-aligned periodic backbone capturing dominant domain morphology; per-variable stochastic residuals modeled with a deep-kernel Gaussian process; and lag-aware cross-variable dependencies injected through a structural causal model with fitted coupling coefficients. Sampling these components at controllable temperature broadens distributional coverage while preserving domain-grounded structure. We show that ReGeN-generated data consistently substitutes for real sibling data with minimal forecasting degradation, and in strongly periodic domains such as traffic, can outperform the real source itself. We further show that a foundation model pretrained on ReGeN corpora outperforms those pretrained on prior-based and data-driven synthetic alternatives. This suggests that in low-data regimes, how reference data is structurally exploited can matter as much as how much data is available.
Moulik Gupta, Dhruv Kumar, Murari Mandal +1
Birla AI Labs, Office of Ananya Birla · Birla Institute of Technology and Science, Pilani · Kalinga Institute of Industrial Technology, Bhubaneswar
Time-series data augmentation plays a crucial role in regression-oriented forecasting tasks, where limited data restricts the performance of deep learning models. While Generative Adversarial Networks (GANs) have shown promise in synthetic time-series generation, existing approaches primarily focus on matching marginal data distributions and often overlook the temporal dynamics that naturally exist in the original multivariate time series. When generating multivariate time series, this mismatch leads to distribution shift and temporal drift, thereby degrading the fidelity of the synthetic sequences. In this work, we propose a model-agnostic Markov Chain Monte Carlo (MCMC)-based framework to mitigate distribution shift and preserve temporal dynamics in synthetic time series. We provide a theoretical analysis of how conditional generative models accumulate deviations under sequential generation and demonstrate that the MCMC algorithm can correct these discrepancies by enforcing consistency with empirical transition statistics between neighboring time points. Extensive experiments on the Lorenz, Licor, ETTh, and ILI datasets using RCGAN, GCWGAN, TimeGAN, SigCWGAN, and AECGAN demonstrate that the proposed MCMC framework consistently improves autocorrelation alignment, skewness error, kurtosis error, R2, discriminative score, and predictive score. These results suggest that synthetic time series consistent with the original data require explicit preservation of transition laws rather than solely relying on adversarial distribution matching, thereby offering a principled direction for improving generative modeling of time-series data.
Weather reanalysis products rely on computationally intensive numerical weather predictions followed by data assimilation that corrects the forecast toward observations. Deep generative models offer a cheaper alternative that shifts much of this cost from inference to offline training. However, existing generative approaches have been evaluated on synthetic observations or under different datasets and evaluation schemes, making it unclear which design choices actually improve real-world data assimilation. We present the first controlled benchmark of generative weather data assimilation on real weather station observations. Using 11,849 NOAA MADIS stations across the contiguous United States and four weather variables, we evaluate methods while holding the dataset, observation operator, and deep learning architecture fixed. The benchmark compares the major design choices, including diffusion versus flow matching, pixel versus latent-space formulations, and multiple inference-time conditioning strategies, against a classical 3D-Var baseline. The benchmark reveals three clear conclusions. First, learned generative priors outperform the Gaussian prior of 3D-Var (35.7% vs. 33.3% RMSE reduction over ERA5) despite using no ERA5 background field at inference. Second, full-gradient guidance consistently outperforms stop-gradient and initial-noise optimization. Third, other choices provide little measurable benefit: diffusion and flow matching perform nearly identically under matched conditions, and latent-space variable mixing does not help. We further evaluate both dense and sparse station settings and find advantages from generative AI and full-gradient guidance more pronounced under sparsity. Together, these results identify which components of generative weather data assimilation improve performance on real station observations and establish a standardized benchmark for future work.
Ruizhe Huang, Qidong Yang, Jonathan Giezendanner +1
Massachusetts Institute of Technology, Cambridge, MA, USA