Synthetic time series are increasingly used for data augmentation, privacy-preserving data sharing, and downstream model development, yet faithfully reproducing both multi-scale temporal structure and cross-channel dependencies remains challenging. We study multivariate time-series generation through flow matching in the wavelet domain. By operating on multilevel discrete wavelet coefficients rather than directly in the time domain, the model represents coarse structure and progressively finer details at separate scales. Their naturally different variances further induce an implicit coarse-to-fine generative process without requiring an explicit multi-scale schedule. Since the transform acts independently on each channel, we pair it with a channel-token transformer whose attention directly models cross-channel dependencies. Across seven benchmark datasets and four sequence lengths, our method is best or tied on a majority of dataset-metric combinations, with the largest and most consistent improvements in Context-FID and discriminative score.
Figures & tables
ETTh1
ETTh2
Stocks
Exchange
EEG
Energy
MuJoCo
Context-FID ( ↓ )
Diffusion-TS
0.137 ± .012
0.056 ± .006
0.175 ± .022
0.047 ± .005
0.018 ± .004
0.090 ± .011
0.016 ± .002
SigDiffusion
2.473 ± .227
1.162 ± .051
3.285 ± .504
1.629 ± .152
0.020 ± .003
4.241 ± .372
2.681 ± .263
FourierDiffusion
0.025 ± .003
0.026 ± .001
0.039 ± .005
0.080 ± .021
0.015 ± .002
0.217 ± .015
0.062 ± .006
WaveletDiff
0.026 ± .002
0.031 ± .002
0.020 ± .002
0.009 ± .000
0.008 ± .001
0.482 ± .042
1.180 ± .095
FlowTS
0.025 ± .001
0.012 ± .001
0.017 ± .005
0.009 ± .001
0.005 ± .000
0.042 ± .004
0.012 ± .001
Table 1: Unconditional generation on short sequences ( T=24 ). Mean ± std over three independent end-to-end training runs with five evaluations each. Bold : best; underlined : second best.
Figure 4: Mean rank of each method over the seven datasets as a function of the window length T (1 = best; tied methods receive the average of their ranks).
Figure 5: t-SNE visualization and probability distributions on Energy. Red is for real data, and blue for generated data.
Figure 6: Effect of removing the implicit coarse-to-fine progression at T=24 . Bars show the ratio of the score under natural coefficient scaling to that under per-level standardization. Values below 1 favor natural scaling and values above 1 favor standardization; the dashed line denotes equal performance.
Variant
ETTh1
ETTh2
Stocks
Exchange
EEG
Energy
MuJoCo
(a) Without ef
1.488 ± 0.099
0.977 ± 0.168
1.763 ± 0.321
0.938 ± 0.072
0.011 ± 0.001
2.617 ± 0.145
1.390 ± 0.139
(b) Std. levels
0.021 ± 0.002
0.014 ± 0.001
0.018 ± 0.004
0.010 ± 0.001
0.010 ± 0.001
0.019 ± 0.002
0.008 ± 0.000
(c) Time domain
0.006 ± 0.001
0.005 ± 0.001
0.014 ± 0.001
0.007 ± 0.000
0.014 ± 0.001
0.012 ± 0.000
0.005 ± 0.000
(d) Uniform t
0.007 ± 0.000
0.008 ± 0.001
0.013 ± 0.001
0.005 ± 0.000
0.014 ± 0.002
0.016 ± 0.001
0.010 ± 0.001
Full model
0.005 ± 0.001
0.004 ± 0.000
0.005 ± 0.002
0.004 ± 0.001
0.013 ± 0.002
0.010 ± 0.003
0.006 ± 0.000
Table 2: Context-FID ( ↓ ) for one-factor-at-a-time ablations at T=24 . The full model uses db2 coefficients, natural level scaling, logit-normal time sampling, and channel-identity embeddings ef . Variants (a)–(d) respectively remove ef , standardize the wavelet levels, replace the wavelet transform with the identity, and sample t uniformly. Results are mean ± standard deviation. Best values are in bold and second-best values are underlined.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Formula
Parameters
Share
MLP, all blocks
B(8dm2+5dm)
4,204,544
42.99%
AdaLN-Zero, all blocks
B(6dm2+6dm)
3,158,016
32.29%
Channel attention, all blocks
B(4dm2+4dm)
2,105,344
21.53%
Time embedding
8dt2+5dt
131,712
1.35%
Output modulation
2dm2+2dm
131,584
1.35%
Conditioning projection
dmdt+dm
33,024
0.34%
Appendix
Table 3: Trainable parameters of the default configuration on ETTh1 at T=24 ( dm=256 , B=8 , dt=128 , MLP ratio 4 , J=3 , D=T=24 coefficients per channel, F=7 ).
Dataset
Type
Length
Channels F
Source
ETTh1
Electricity transformer
17,420
7
Zhou et al. [2021]
ETTh2
Electricity transformer
17,420
7
Zhou et al. [2021]
Stocks
Daily Google prices
3,685
6
Yoon et al. [2019]
Exchange
Daily exchange rates
7,588
8
Lai et al. [2018]
EEG
14-electrode recording
14,980
14
Roesler [2013]
Energy
Appliance energy
19,735
28
Candanedo et al. [2017]
Appendix
Table 4: Datasets used in our experiments. Length is the number of time steps in the raw series; for MuJoCo, 10,000 trajectories of length T are simulated.
Model
Sampler
NFE per sample
Diffusion-TS
DDPM / DDIM
100–1000 †
SigDiffusion
Tsit5 (128 steps)
768
FourierDiffusion
VP-SDE (1000 steps)
1000
WaveletDiff
DDPM (1000 steps)
1000
FlowTS
Euler (100 steps)
100
Ours
Euler (100 steps)
100
Appendix
Table 5: Network function evaluations (NFE) per generated sample at T=24 . † Diffusion-TS uses 500 evaluations on ETTh1, ETTh2, Stocks, and Exchange, 100 on EEG, and 1000 on Energy and MuJoCo.
a3
d3
d2
d1
Dataset
σ
t
σ
t
σ
t
σ
t
ETTh1
2.57
0.28
0.90
0.53
0.42
0.70
0.24
0.81
ETTh2
2.76
0.27
0.40
0.71
0.24
0.81
0.17
0.85
Stocks
2.75
0.27
0.30
0.77
0.23
0.81
0.17
0.85
Exchange
2.82
0.26
0.12
0.89
0.07
0.93
0.04
0.96
EEG
1.41
0.41
0.95
0.51
0.93
0.52
0.92
0.52
Appendix
Table 6: Per-level coefficient standard deviation σ and SNR crossing time t=1/(1+σ) at T=24 , under the db2 wavelet. Coarser levels have larger σ and therefore cross earlier.
Figure 7: Level-wise SNR along the linear probability path on the other six datasets. Coarse levels cross SNR=1 earlier than fine levels.
ETTh1
ETTh2
Stocks
Exchange
EEG
Energy
MuJoCo
Context-FID ( ↓ )
Diffusion-TS
0.187 ± .006
0.073 ± .004
0.238 ± .050
0.048 ± .005
0.024 ± .002
0.106 ± .019
0.021 ± .002
SigDiffusion
2.972 ± .112
1.413 ± .181
4.086 ± .784
1.702 ± .101
0.032 ± .004
4.730 ± .547
3.117 ± .150
FourierDiffusion
0.039 ± .005
0.034 ± .004
0.055 ± .010
0.058 ± .009
0.022 ± .001
0.296 ± .018
0.099 ± .012
WaveletDiff
0.059 ± .005
0.074 ± .005
0.024 ± .002
0.007 ± .001
0.011 ± .002
0.541 ± .022
1.269 ± .149
FlowTS
0.030 ± .002
0.013 ± .001
0.019 ± .003
0.010 ± .001
0.008 ± .001
0.061 ± .004
0.026 ± .001
Appendix
Table 7: Unconditional generation at T=32 . Mean ± std over five recomputations of each metric on the same generated set. Bold : best; underlined : second best.
ETTh1
ETTh2
Stocks
Exchange
EEG
Energy
MuJoCo
Context-FID ( ↓ )
Diffusion-TS
0.267 ± .023
0.112 ± .010
0.354 ± .080
0.056 ± .003
0.058 ± .005
0.085 ± .007
0.035 ± .002
SigDiffusion
5.948 ± .465
1.581 ± .174
3.851 ± .744
1.986 ± .125
0.056 ± .004
6.403 ± .271
3.621 ± .181
FourierDiffusion
0.089 ± .005
0.068 ± .005
0.111 ± .012
0.081 ± .008
0.045 ± .005
0.446 ± .030
0.167 ± .009
WaveletDiff
0.104 ± .008
0.072 ± .003
0.059 ± .011
0.151 ± .015
0.025 ± .003
0.438 ± .013
0.347 ± .021
FlowTS
0.052 ± .003
0.027 ± .002
0.038 ± .005
0.016 ± .001
0.015 ± .003
0.172 ± .016
0.093 ± .005
Appendix
Table 8: Unconditional generation at T=64 . Same conventions as .
ETTh1
ETTh2
Stocks
Exchange
EEG
Energy
MuJoCo
Context-FID ( ↓ )
Diffusion-TS
0.994 ± .025
0.146 ± .018
0.611 ± .068
0.059 ± .004
0.139 ± .023
0.074 ± .008
0.073 ± .004
SigDiffusion
11.619 ± .739
2.087 ± .284
4.743 ± .381
2.005 ± .267
0.130 ± .012
10.471 ± .272
3.738 ± .269
FourierDiffusion
0.319 ± .020
0.200 ± .013
0.272 ± .023
0.235 ± .040
0.086 ± .014
0.837 ± .034
0.224 ± .021
WaveletDiff
0.158 ± .023
0.088 ± .006
0.112 ± .010
0.154 ± .015
0.059 ± .005
0.501 ± .037
0.277 ± .018
FlowTS
0.097 ± .003
0.058 ± .004
0.063 ± .010
0.035 ± .004
0.027 ± .003
0.260 ± .021
0.156 ± .009
Appendix
Table 9: Unconditional generation at T=128 . Same conventions as .
Figure 8: Evolution of each metric with the window length T .
Figure 9: t-SNE embeddings of real (red) and generated (blue) windows for all methods and datasets at T=24 .
Figure 10: Probability distributions of real (red) and generated (blue) data for all methods and datasets at T=24 .
Variant
ETTh1
ETTh2
Stocks
Exchange
EEG
Energy
MuJoCo
Context-FID ( ↓ )
(a) Without ef
1.488 ± 0.099
0.977 ± 0.168
1.763 ± 0.321
0.938 ± 0.072
0.011 ± 0.001
2.617 ± 0.145
1.390 ± 0.139
(b) Std. levels
0.021 ± 0.002
0.014 ± 0.001
0.018 ± 0.004
0.010 ± 0.001
0.010 ± 0.001
0.019 ± 0.002
0.008 ± 0.000
(c) Time domain
0.006 ± 0.001
0.005 ± 0.001
0.014 ± 0.001
0.007 ± 0.000
0.014 ± 0.001
0.012 ± 0.000
0.005 ± 0.000
(d) Uniform t
0.007 ± 0.000
0.008 ± 0.001
0.013 ± 0.001
0.005 ± 0.000
0.014 ± 0.002
0.016 ± 0.001
0.010 ± 0.001
Full model
0.005 ± 0.001
0.004 ± 0.000
0.005 ± 0.002
0.004 ± 0.001
0.013 ± 0.002
0.010 ± 0.003
0.006 ± 0.000
Appendix
Table 10: One-factor-at-a-time ablations at T=24 . The full model uses db2 coefficients, natural level scaling, logit-normal time sampling, and channel-identity embeddings ef . Variants (a)–(d) respectively remove ef , standardize the wavelet levels, replace the wavelet transform with the identity, and sample t uniformly. Results are mean ± standard deviation; Best values are in bold and second-best values are underlined. The shaded rows show the full model.
Time shift α
Dataset
0.33
0.5
1.0†
2.0
3.0
5.0
Context-FID ( ↓ )
ETTh1
0.007 ± 0.000
0.008 ± 0.000
0.005 ± 0.001
0.011 ± 0.001
0.010 ± 0.000
0.021 ± 0.001
ETTh2
0.008 ± 0.000
0.005 ± 0.000
0.004 ± 0.000
0.004 ± 0.000
0.005 ± 0.000
0.007 ± 0.000
Stocks
0.008 ± 0.001
0.011 ± 0.001
0.005 ± 0.002
0.006 ± 0.001
0.010 ± 0.001
0.023 ± 0.003
Exchange
0.005 ± 0.001
0.005 ± 0.000
0.004 ± 0.001
0.003 ± 0.000
0.007 ± 0.000
0.008 ± 0.001
Appendix
Table 11: Effect of the sampling time shift α at T=24 , with N=100 Euler steps, all from the same trained model. † Default setting. Results are mean ± std; the best value in each row is in bold.
Euler steps N
Dataset
25
50
100†
200
400
Context-FID ( ↓ )
ETTh1
0.019 ± 0.002
0.011 ± 0.001
0.005 ± 0.001
0.007 ± 0.000
0.008 ± 0.001
ETTh2
0.009 ± 0.001
0.005 ± 0.001
0.004 ± 0.000
0.005 ± 0.000
0.009 ± 0.000
Stocks
0.021 ± 0.003
0.004 ± 0.001
0.005 ± 0.002
0.003 ± 0.000
0.010 ± 0.001
Exchange
0.009 ± 0.000
0.003 ± 0.000
0.004 ± 0.001
0.005 ± 0.000
0.006 ± 0.000
Appendix
Table 12: Effect of the number of Euler steps N at T=24 , with time shift α=1 , all from the same trained model. † Default setting. Results are mean ± std; the best value in each row is in bold.
Transform
ETTh1
ETTh2
Stocks
Exchange
EEG
Energy
MuJoCo
Context-FID ( ↓ )
db2
0.005 ± 0.001
0.004 ± 0.000
0.005 ± 0.002
0.004 ± 0.001
0.013 ± 0.002
0.010 ± 0.003
0.006 ± 0.000
coif1
0.005 ± 0.000
0.003 ± 0.000
0.009 ± 0.001
0.009 ± 0.001
0.014 ± 0.001
0.008 ± 0.000
0.007 ± 0.001
bior2.2
0.006 ± 0.000
0.005 ± 0.000
0.006 ± 0.000
0.004 ± 0.001
0.014 ± 0.001
0.013 ± 0.001
0.006 ± 0.001
rbio2.2
0.004 ± 0.000
0.005 ± 0.000
0.019 ± 0.001
0.006 ± 0.001
0.010 ± 0.002
0.013 ± 0.001
0.006 ± 0.000
Learned
0.004 ± 0.000
0.003 ± 0.000
0.008 ± 0.001
0.008 ± 0.000
0.011 ± 0.002
0.009 ± 0.001
0.005 ± 0.000
Appendix
Table 13: Effect of the wavelet transform at T=24 . Results are mean ± std; the best transform for each dataset and metric is in bold.
Figure 11: Deviation of the lattice angles from their db2 initialization during training at T=24 .
Vector quantization (VQ) with autoregressive (AR) token modeling is a widely adopted and highly competitive paradigm for time-series generation. However, such models are fundamentally limited by exposure bias: during inference, errors can accumulate across sequential predictions, leading to pronounced quality degradation in long-horizon generation. To address this, we propose SDFlow (Similarity-Driven Flow Matching), a non-autoregressive framework that operates entirely in the frozen VQ latent space and enables parallel sequence generation via flow matching. We tackle three key challenges in making this transition: (1) eliminating exposure bias by replacing step-wise token prediction with a global transport map; (2) mitigating the high-dimensionality of VQ token spaces via a low-rank manifold decomposition with a learned anchor prior over the latent manifold; and (3) incorporating discrete supervision into continuous transport dynamics by introducing a categorical posterior over codebook indices within a variational flow-matching formulation. Extensive experiments show that SDFlow achieves state-of-the-art performance, improving Discriminative Score and substantially reducing Context-FID, particularly for challenging long-sequence generation. Moreover, SDFlow provides significant inference speedups over autoregressive baselines, offering both high fidelity and computational efficiency. Code is available at https://anonymous.4open.science/r/SDFlow-D6F3/
Wei Li, Shibo Feng, Pengcheng Wu +3
1Shanghai Jiao Tong University · 2Shanghai University · 3Nanyang Technological University +2
Generating high-quality time-series data is challenging because real-world signals often exhibit multimodal patterns and multiscale dynamics, including oscillations and high-frequency variations. Flow Matching (FM) offers an efficient alternative to diffusion models, but practical implementations typically rely on a single finite-capacity global vector-field estimator. In such heterogeneous temporal distributions, distinct regimes may pass through nearby flow states while requiring incompatible conditional velocities. A monolithic estimator trained with the standard ℓ2 velocity-matching objective may therefore learn an overly smoothed approximation of the local transport field. This estimator-level smoothing can attenuate branch-specific dynamics, leading to spectral distortion and poor mode coverage. To address this, we propose PrismFlow, a new FM method with Koopman-inspired dynamical experts. Each expert learns residual corrections in a latent space where local nonlinear temporal evolution can be approximated by linear transitions. We further propose a confidence-aware Winner-Take-All (WTA) objective that updates only the expert best aligned with each sample while masking gradients to the others, encouraging mode-specific specialization. During sampling, the selected expert adds a residual dynamical correction to the global transport field, preserving FM stability while recovering fine-grained and high-frequency temporal structures. Across various benchmarks, PrismFlow effectively mitigates the spectral contraction in standard FM and achieves state-of-the-art performance, with a 15.6% gain in Context-FID and a 38.6% improvement in Discriminative Score, while remaining robust in low-data settings and effective for forecasting and imputation.
Junru Zhang, Lang Feng, Jinbo Wang +6
Zhejiang University, China · Nanyang Technological University, Singapore · I2R, Agency for Science, Technology and Research (A*STAR), Singapore
As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for generation quality. However, standard flow matching treats all spatial frequencies uniformly, ignoring the natural frequency hierarchy where high-frequency bands become indistinguishable from pure noise far earlier than coarse structures. We introduce WaiT, a Wavelet-aware image Transformer that decomposes generation into coarse and fine bands via lossless wavelets. True to its name, the high-frequency bands wait for the signal: staying pure noise until coarse structure has emerged, then joining the flow for joint refinement. Since standard FID discards fine-grained detail through aggressive downsampling, we introduce a more stringent three-axis evaluation protocol to assess quality at native resolution. On ImageNet 512x512, WaiT achieves a pixel-space FID of 1.43 and is Pareto-optimal across all three axes, reducing sampling compute by up to 50%. With our largest 2B model, we set a new state-of-the-art FID of 1.3 for pixel-space models on ImageNet 512 resolution. Our formulation outperforms even the strongest latent-space models on texture fidelity, and scales seamlessly to high-resolution OpenImages and to video generation, achieving a state-of-the-art FVD of 0.84 on Kinetics-600 with no algorithmic modifications.