cs.LGSep 27, 2026

StatD2GAN: When Calibration Masks Generator Quality in Held-Out Evaluation of Synthetic Weather Sequences

Authors: Mustafa Ozaytac, Ozge Karadag Atas

Organizations: Department of Statistics, Graduate School of Science and Engineering, Hacettepe University, Ankara, 06800, Turkey · Department of Statistics, Faculty of Science, Hacettepe University, Ankara, 06800, Turkey

Abstract

Generative models for multivariate weather series are routinely evaluated with pooled distributional metrics computed after marginal calibration. We show this practice can invalidate architectural conclusions, and rebuild the evaluation of StatD2GAN, a three-discriminator GAN with evolutionary weight adaptation, around a held-out protocol: the final two calendar years of each dataset are held out behind a 168 hour embargo, calibration is fitted on the training block only, and all metrics are computed on the held-out block. Evidence comes from 25 matched (location, seed) pairs across five Koppen-Geiger climates, tested with Wilcoxon signed-rank tests under Holm correction. Four results follow. First, isotonic calibration drives the Kolmogorov-Smirnov distance to within 2% of a per-location noise-and-shift floor for every architecture tested, including a deliberately weak RCGAN baseline, so calibrated marginal metrics cannot discriminate between architectures. Second, the sorted-representation discriminator is the only component whose removal significantly degrades cross-variable dependence (Kendall tau MAE +0.080, Holm p = 0.009), with a regime-dependent effect: near zero in Ankara, above 115% in Dubai and Yakutsk. A rank-transformed variant isolates the mechanism as quantile supervision of the marginals rather than copula matching. Third, physical constraint violations are injected by calibration, not the generator; projection removes them at negligible cost (deltaKS <= 0.003). Fourth, pooled metrics conceal a collapse of between-sequence weekly-mean variability, a proxy for seasonal and regime diversity, in TimeGAN that only sequence-level statistics expose. We recommend floor-referenced marginal evaluation, matched-pair testing, and sequence-level variance decomposition as minimum requirements for calibrated generative pipelines.

Figures & tables

Explore similar work

CardsList
  1. REGEN: Reference-Guided Synthetic Multivariate Time Series Generation for Forecasting

    Jun 3, 2026Moulik Gupta, Dhruv Kumar, Murari Mandal +1Time-Series GenerationMultivariate Time Series Forecasting

  2. Benchmarking Generative Models for Weather Data Assimilation on Real Station Observations

    Sep 30, 2026Ruizhe Huang, Qidong Yang, Jonathan Giezendanner +1Data AssimilationEra5 Reanalysis Data