Differentially private time-series generators commonly produce fixed-length synthetic windows, whereas downstream forecasting models often require long continuous training sequences. How these windows are assembled after generation can therefore alter the effective synthetic data presented to a forecaster, even when the trained generator remains unchanged. We study this post-generation sequence assembly process by systematically varying overlap rates and window-weighting schemes and evaluating the resulting sequences in terms of boundary continuity, statistical and temporal fidelity, and Train-on-Synthetic-Test-on-Real (TSTR) forecasting utility. Across four types of public datasets (ETTh1, ETTm1, Weather, and Appliances) and five forecasting models, the results reveal a clear forecaster-dependent assembly principle: downstream TSTR utility is jointly shaped by the forecaster, overlap rate, and window-weighting scheme, leading to distinct assembly preferences across forecasting models. Increased overlap generally improves boundary continuity, but improvements in continuity or individual fidelity diagnostics do not consistently reduce forecasting error, indicating that these diagnostics alone are insufficient for selecting assembly configurations. Complete five-forecaster assembly grids, together with matched Train-on-Real-Test-on-Real (TRTR) references, further characterize these regularities and quantify assembly-dependent utility relative to real-data training. We then validate the identified principles through additional analyses of robustness and generator variability.
Figures & tables
Figure 1: Illustration of post-generation sequence assembly on ETTh1. Left: direct concatenation of independently generated windows creates a visible boundary discontinuity. Right: forecasting performance of the real-data training reference (blue) under Train-on-Real-Test-on-Real (TRTR) evaluation and the naive-concatenation synthetic-data baseline (red) under Train-on-Synthetic-Test-on-Real (TSTR) evaluation.
Dataset
Naive Jjump
Evaluated assembly Jjump
Reduction
ETTh1
4.7745
0.3964
91.7%
ETTm1
4.1697
0.1412
96.6%
Weather
47.711
29.853
37.4%
Appliances
216.001
94.367
56.3%
Table 2: Boundary discontinuity reduction across datasets. Values are seed-averaged means; the assembly column reports the lowest Jjump among the evaluated high-overlap settings.
Figure 2: ETTh1 assembly path under triangular weighting. Boundary discontinuity decreases while the physical 24-hour autocorrelation deviation changes; LSTM and ARIMA TSTR MAE show different downstream responses. The corresponding endpoint comparisons across the four public evaluation datasets are reported in Table 3 .
Figure 3: Minimum-MAE assembly configurations identified from the seed-averaged grids for the four public evaluation datasets. Each cell represents one dataset–forecaster pair. Color intensity represents overlap rate ρ ; text labels denote Tri (Triangular weighting), Hann (Hann weighting), and None (no explicit windowing).
Forecaster
TRTR MAE
Naive TSTR MAE
TSTR range
LSTM
0.955
1.685
1.147–1.724
RNN
1.306
1.961
1.640–2.645
LightGBM
0.882
0.900
0.900–1.435
ARIMA
1.542
2.667
2.599–3.092
Prophet
1.399
2.220
2.220–3.148
Table 5: Operational electricity-load forecasting results. All entries are MAE; the TSTR range spans the evaluated assembly settings.
Dataset
Jjump
LSTM MAE
ARIMA MAE
ETTh1
0.6892 → 0.1233 → 0.0549
3.080 → 2.947 → 2.814
2.952 → 3.215 → 4.154
ETTm1
0.4812 → 0.2125 → 0.1770
2.848 → 2.822 → 2.776
4.020 → 4.225 → 4.319
Table 6: DP-ConvVAE results across three assembly settings. Values are means; entries are ordered C0 → C1 → C2.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Backbone
Conditional 1D U-Net
Input channels
1
Output channels
1
Base model channels
144
Channel multipliers
1,2,3,4
Number of residual blocks
2
Appendix
Table A.3: Diffusion architecture and DP-SGD training hyperparameters.
Dataset
Noise multiplier σ
ETTh1
2.019
ETTm1
2.078
Weather
2.024
Appliances
2.024
Operational load
2.041
Appendix
Table B.1: Dataset-specific DP-SGD noise multipliers used for generator training. The values are calibrated using Opacus to satisfy the target (ϵtrain,δtrain) guarantee under the corresponding training schedule.
Figure C.1: Additional diagnostic relationship analysis. Left: relationship between boundary discontinuity ( Jjump ) and TSTR MAE for LSTM. Right: relationship between 24-hour autocorrelation deviation ( ΔACF24h ) and TSTR MAE for ARIMA. Each point corresponds to one overlap configuration in the evaluated assembly path.
Time series forecasting in privacy-sensitive domains often requires training models on released data rather than original observations. Synthetic time series generation has been developed primarily for data augmentation, where generated series supplement the original training set. How well these methods perform when fully replacing the original data - and how much privacy risk the released series carry - remains underexplored. We address this gap through a benchmark evaluating synthetic generation methods and noise-based anonymization baselines under a Train on Synthetic, Test on Real (TSTR) protocol. We jointly assess forecasting performance and distance-based empirical privacy risk across seven datasets, characterizing the trade-off between these objectives. We also introduce Grasynda-P, a privacy-motivated extension of the graph-based generator Grasynda, incorporating matrix ensembling and kernel density estimation. Our results show that: (1) no generation method fully substitutes for original training data; (2) noise-based anonymization yields the strongest privacy but the worst forecasting performance; (3) simple transformation-based generators outperform deep generative models for forecasting in this setting; and (4) Grasynda-P lies on the Pareto frontier, achieving competitive forecasting with stronger privacy separation than other generators. This benchmark establishes a reference point for evaluating and developing new privacy-aware synthetic time series generation methods.
Luis Amorim, Vitor Cerqueira, Moises Santos +2
Faculdade de Engenharia da Universidade do Porto, Porto, Portugal · University of Coimbra, Coimbra, Portugal · Laboratory for Artificial Intelligence and Computer Science (LIACC), Portugal +3
Choosing the wrong synthetic generator for time-series foundation model pretraining is costly: under identical training budgets, the best and worst generators produce up to a 2× gap in forecasting error, yet the field has no principled way to make this choice. The problem is compounded by the fact that generator rankings are not stable across architectures: across 11 generator families evaluated on Chronos-T5-Mini and Moirai-Small trained from scratch, we find that which generators are useful depends on the model architecture. Rather than solving the generator selection problem, we sidestep it: a simple equal-weight mixture of all generators matches or beats the best individual generator for both architectures, and composing this mixture with real data yields the strongest pretraining corpora overall. Synthetic pretraining is therefore a corpus composition problem, not a generator selection problem, and composition choices should be validated per model family rather than assumed to transfer.
Aaryan Nagpal, Debdeep Sanyal, Murari Mandal +2
Birla AI Labs, Mumbai, India · KIIT, Bhubaneswar, India · BITS Pilani, Pilani, India
Synthetic data has transformed language model training, yet its role in time series forecasting remains poorly understood. We present a large-scale empirical study: nine experiment groups, 4,218 runs systematically evaluating synthetic time series augmentation across five architectures, four synthetic signals and seven datasets. The effect is sharply architecture-conditional: channel-mixing models (TimesNet, iTransformer) benefit in the majority of trials, while channel-independent models (DLinear, PatchTST) are consistently degraded. In selected low-resource settings the gains are striking: TimesNet trained on only 10% of Weather data with synthetic augmentation surpasses the full-data baseline (4 of 16 sparsity-dataset combinations). Averaged across all architectures, augmentation hurts in 67% of trials. We further find that only the Seasonal-Trend generator reliably helps across the tested benchmarks, and that hard curriculum switching is actively harmful (+24% MSE degradation). These results provide concrete, actionable guidelines on how to use synthetic data: use synthetic augmentation with channel-mixing architectures, use gradual annealing schedules, and treat low-resource augmentation as architecture- and dataset-dependent. Code is available at \href{https://github.com/hugoiscracked/synthetic-ts/tree/main}
Hugo Cazaux, Eyjólfur Ingi Ásgeirsson, Hlynur Stefánsson
Department of Engineering Reykjavík University Iceland, Menntavegur 1, 102