Generative modeling has shown strong promise for multivariate time mseries (MTS) forecasting, especially scale to high-dimensional settings. Diffusion-based methods achieve competitive performance but typically require many sampling steps at inference. VAE-based non-iterative forecasting frameworks have therefore emerged as an efficient alternative. Within this line of work, vector quantization (VQ) enables controllable latent space modeling by mapping multivariate series into compact discrete representations. Existing VQ-based forecasting methods, however, typically rely on autoregressive (AR) token generation, which suffers from exposure bias and training-inference mismatch. Flow matching provides an efficient non-autoregressive alternative for latent forecasting, but existing formulations usually initialize transport from a generic Gaussian prior. We instead observe that the trained VQ codebook already captures representative latent prototypes and can thus serve as a more informative prior for flow matching. Based on this insight, we propose ProtoFlow, a forecasting framework that combines vector-quantized autoencoding with Prototype-prior Flow matching. Our method first maps multivariate sequences into a discrete latent space, then constructs a structured prior from the learned codebook, and finally learns a DiT-based rectified flow to transport samples from this prior to future latent representations conditioned on historical observations. By replacing generic noise initialization with a learned prototype prior, ProtoFlow avoids the rollout mismatch of AR token prediction and promotes faster training convergence. Extensive experiments on benchmark datasets show that it consistently achieves superior forecasting performance with efficient inference.
Figures & tables
Figure 1: Illustration of our ProtoFlow. Stage 1 discretizes the future target sequence into compact latent tokens via similarity-based VQ latent modeling. Stage 2 compares three latent-generation paradigms under the same tokenized space: (a) Autoregressive rollout, which suffers from training–inference mismatch; (b) Gaussian-initialized flow matching, which starts transport from an unstructured prior; and (c) The proposed ProtoFlow, which initializes conditional rectified flow from a learned codebook prior for latent flow matching generation. The future time features yc is omitted here.
Electricity
Solar
Traffic
Taxi
Wikipedia
Horizon+Method
CRPSsum
NRMSEsum
CRPSsum
NRMSEsum
CRPSsum
NRMSEsum
CRPSsum
NRMSEsum
CRPSsum
NRMSEsum
48 Steps
∙ TimeGrad ( 2021 )
0.031
0.044
0.324
0.736
0.043
0.072
0.252
0.334
0.082
0.105
∙ CSDI ( 2021 )
0.029
0.050
0.328
0.742
-
-
-
-
-
-
∙ TSDiff ( 2023 )
0.036
0.057
0.324
0.731
0.057
0.070
0.243
0.330
0.079
0.091
∙ MG-TSD ( 2024 )
0.024
0.038
0.328
0.737
0.042
0.067
0.217
0.327
0.071
0.096
Table 1: The detailed results of ProtoFlow on the 48 and 96 prediction horizons. ProtoFlow-G: the variant of ProtoFlow that uses a standard Gaussian prior. - marks out-of-memory failures, Best and second-best forecasting results are bolded and underlined, respectively.
Method
Electricity
Solar
Traffic
Taxi
Wikipedia
Avg.
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
VAE + AR
0.421
0.544
1.154
2.618
1.005
1.511
0.337
0.448
0.130
0.169
0.609
1.058
VAE + FM
0.129
0.213
0.643
1.324
0.323
0.564
0.227
0.293
0.117
0.128
0.288
0.504
VQ + AR
0.037
0.054
0.347
0.833
0.066
0.084
0.205
0.322
0.108
0.121
0.153
0.283
ProtoFlow-G
0.019
0.037
0.305
0.717
0.036
0.058
0.159
0.287
0.096
0.114
0.123
0.243
ProtoFlow
0.015
0.022
0.284
0.638
0.034
0.054
0.157
0.281
0.092
0.101
0.116
0.219
Table 2: The comparison results of 96 prediction horizons with baselines regarding CRPSsum (C-S) and NRMSEsum (NM-S). AR: Auto-regressive Transformer and FM: Flow Matching .
Figure 2: Comparison of training convergence ( a ), inference efficiency, training memory ( b ), and forecasting performance ( c-d ) on Electricity, Traffic, and Solar with a prediction length of 96.
Method
Electricity
Solar
Traffic
Taxi
Wikipedia
Avg.
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
ProtoFlow-G
0.019
0.037
0.305
0.717
0.036
0.058
0.159
0.287
0.096
0.114
0.123
0.243
ProtoFlow-uniform
0.018
0.030
0.356
0.769
0.036
0.056
0.157
0.284
0.099
0.108
0.133
0.249
ProtoFlow
0.015
0.022
0.284
0.638
0.034
0.054
0.157
0.281
0.092
0.101
0.116
0.219
Table 3: The forecasting performance with frequency-weighted prototype (ProtoFlow), uniform-weight prototype (ProtoFlow-uniform), and isotropic Gaussian priors (ProtoFlow-G) at a prediction length of 96.
Figure 3: Effects of codebook size on ProtoFlow at the 96-step prediction length.
Electricity
Solar
Traffic
Frequency
Uniform
Stage1 MSE
Frequency
Uniform
Stage1 MSE
Frequency
Uniform
Stage1 MSE
Coodebook size
64
0.039
0.044
7.76e-4
1.035
0.962
7.70e-4
0.113
0.105
5.82e-4
128
0.030
0.032
7.13e-4
0.774
0.761
6.75e-4
0.062
0.066
5.23e-4
256
0.022
0.030
6.02e-4
0.638
0.769
6.13e-4
0.054
0.056
4.78e-4
512
0.031
0.033
6.01e-4
0.993
0.758
6.18e-4
0.056
0.057
4.69e-4
Table 4: Effects of codebook size on Stage 1 reconstruction MSE and NRMSEsum under frequency and uniform weighting at a 96-step prediction length.
Figure 4: Effect of ODE sampling steps on 96-step probabilistic forecasting performance and W2 transport distance for ProtoFlow and ProtoFlow-G.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
DATASET
Dimension
Domain
Freq
Total Time Steps
Pred Length
Rolling Windows
SOLAR
137
R+
Hourly
7,009
{48,96}
7
ELECTRICITY
370
R+
Hourly
5,790
{48,96}
7
TRAFFIC
963
(0,1)
Hourly
10,413
{48,96}
7
TAXI
1214
N
30-Min
1,488
{48,96}
57
WIKIPEDIA
2000
N
Daily
792
{48,96}
5
Appendix
Table 5: The dataset properties used in our experiments are summarized as follows.
Datasets
Codebook Size
Codebook dim
Enc/Dec Layers
λ
β
Solar
256
128
2
0.99
0.1
Electricity
256
256
2
0.99
0.1
Traffic
256
256
2
0.99
0.1
Taxi
256
512
2
0.99
0.05
Wikipedia
256
512
2
0.99
0.05
Appendix
Table 6: Detailed hyperparameters of Stage 1.
Datasets
Hidden dim
DiTθ layers
Batch size
Steps
lr
Solar
128
4
64
3
5e-4
Electricity
256
4
64
3
5e-4
Traffic
256
5
64
3
3e-4
Taxi
512
5
64
3
3e-4
Wikipedia
512
4
64
5
3e-4
Appendix
Table 7: Detailed hyperparameters of Stage 2.
Layer
Function
Descriptions
1
Convolution
input channel= d , output channel=D, kernel size=3, stride=1, padding=1
Figure 6: ProtoFlow forecasts with a prediction horizon of 96. Rows show Electricity (370 dimensions), Traffic (963 dimensions), Solar (137 dimensions), Taxi (1214 dimensions), and Wikipedia (2000 dimensions) from top to bottom, with four examples per dataset.
Vector quantization (VQ) with autoregressive (AR) token modeling is a widely adopted and highly competitive paradigm for time-series generation. However, such models are fundamentally limited by exposure bias: during inference, errors can accumulate across sequential predictions, leading to pronounced quality degradation in long-horizon generation. To address this, we propose SDFlow (Similarity-Driven Flow Matching), a non-autoregressive framework that operates entirely in the frozen VQ latent space and enables parallel sequence generation via flow matching. We tackle three key challenges in making this transition: (1) eliminating exposure bias by replacing step-wise token prediction with a global transport map; (2) mitigating the high-dimensionality of VQ token spaces via a low-rank manifold decomposition with a learned anchor prior over the latent manifold; and (3) incorporating discrete supervision into continuous transport dynamics by introducing a categorical posterior over codebook indices within a variational flow-matching formulation. Extensive experiments show that SDFlow achieves state-of-the-art performance, improving Discriminative Score and substantially reducing Context-FID, particularly for challenging long-sequence generation. Moreover, SDFlow provides significant inference speedups over autoregressive baselines, offering both high fidelity and computational efficiency. Code is available at https://anonymous.4open.science/r/SDFlow-D6F3/
Wei Li, Shibo Feng, Pengcheng Wu +3
1Shanghai Jiao Tong University · 2Shanghai University · 3Nanyang Technological University +2
Generating high-quality time-series data is challenging because real-world signals often exhibit multimodal patterns and multiscale dynamics, including oscillations and high-frequency variations. Flow Matching (FM) offers an efficient alternative to diffusion models, but practical implementations typically rely on a single finite-capacity global vector-field estimator. In such heterogeneous temporal distributions, distinct regimes may pass through nearby flow states while requiring incompatible conditional velocities. A monolithic estimator trained with the standard ℓ2 velocity-matching objective may therefore learn an overly smoothed approximation of the local transport field. This estimator-level smoothing can attenuate branch-specific dynamics, leading to spectral distortion and poor mode coverage. To address this, we propose PrismFlow, a new FM method with Koopman-inspired dynamical experts. Each expert learns residual corrections in a latent space where local nonlinear temporal evolution can be approximated by linear transitions. We further propose a confidence-aware Winner-Take-All (WTA) objective that updates only the expert best aligned with each sample while masking gradients to the others, encouraging mode-specific specialization. During sampling, the selected expert adds a residual dynamical correction to the global transport field, preserving FM stability while recovering fine-grained and high-frequency temporal structures. Across various benchmarks, PrismFlow effectively mitigates the spectral contraction in standard FM and achieves state-of-the-art performance, with a 15.6% gain in Context-FID and a 38.6% improvement in Discriminative Score, while remaining robust in low-data settings and effective for forecasting and imputation.
Junru Zhang, Lang Feng, Jinbo Wang +6
Zhejiang University, China · Nanyang Technological University, Singapore · I2R, Agency for Science, Technology and Research (A*STAR), Singapore
Probabilistic forecasting is important for predicting complex dynamical systems because intrinsic randomness and incomplete observations can cause the same observed state to evolve into multiple plausible futures. While flow matching is a flexible approach for probabilistic forecasting, it is computationally expensive. Streaming flow (SF) reformulates this approach to model temporal evolution efficiently by learning a continuous velocity field directly in physical time. However, SF learns a deterministic velocity field. Thus, it provides only a single future trajectory for a given fixed initial state and observation history. To overcome this limitation, we introduce Variational Streaming Flow (VSF). Our approach learns a latent distribution that is conditioned on the dynamics of interest. In turn, this enables probabilistic forecasting. Importantly, we retain the computational efficiency of SF by generating in physical time. Across deterministic and stochastic dynamical systems, VSF demonstrates superior predictive accuracy and distributional fidelity. We demonstrate the advantage for both long-horizon rollouts exceeding 1,000 steps, and settings with bifurcating dynamics. Moreover, VSF can be integrated into existing Joint-Embedding Predictive Architecture (JEPA)-based world models as a plug-and-play predictor to improve temporal dynamics and goal-directed success rate in navigation, motion planning, and manipulation.
Hans Hao-Hsun Hsu, Minseon Gwak, Soon Hoe Lim +2
Georgia Institute of Technology · Lawrence Berkeley National Lab · International Computer Science Institute +2