Generative modeling has shown strong promise for multivariate time mseries (MTS) forecasting, especially scale to high-dimensional settings. Diffusion-based methods achieve competitive performance but typically require many sampling steps at inference. VAE-based non-iterative forecasting frameworks have therefore emerged as an efficient alternative. Within this line of work, vector quantization (VQ) enables controllable latent space modeling by mapping multivariate series into compact discrete representations. Existing VQ-based forecasting methods, however, typically rely on autoregressive (AR) token generation, which suffers from exposure bias and training-inference mismatch. Flow matching provides an efficient non-autoregressive alternative for latent forecasting, but existing formulations usually initialize transport from a generic Gaussian prior. We instead observe that the trained VQ codebook already captures representative latent prototypes and can thus serve as a more informative prior for flow matching. Based on this insight, we propose ProtoFlow, a forecasting framework that combines vector-quantized autoencoding with Prototype-prior Flow matching. Our method first maps multivariate sequences into a discrete latent space, then constructs a structured prior from the learned codebook, and finally learns a DiT-based rectified flow to transport samples from this prior to future latent representations conditioned on historical observations. By replacing generic noise initialization with a learned prototype prior, ProtoFlow avoids the rollout mismatch of AR token prediction and promotes faster training convergence. Extensive experiments on benchmark datasets show that it consistently achieves superior forecasting performance with efficient inference.
Figures & tables
Figure 1: Illustration of our ProtoFlow. Stage 1 discretizes the future target sequence into compact latent tokens via similarity-based VQ latent modeling. Stage 2 compares three latent-generation paradigms under the same tokenized space: (a) Autoregressive rollout, which suffers from training–inference mismatch; (b) Gaussian-initialized flow matching, which starts transport from an unstructured prior; and (c) The proposed ProtoFlow, which initializes conditional rectified flow from a learned codebook prior for latent flow matching generation. The future time features yc is omitted here.
Electricity
Solar
Traffic
Taxi
Wikipedia
Horizon+Method
CRPSsum
NRMSEsum
CRPSsum
NRMSEsum
CRPSsum
NRMSEsum
CRPSsum
NRMSEsum
CRPSsum
NRMSEsum
48 Steps
∙ TimeGrad ( 2021 )
0.031
0.044
0.324
0.736
0.043
0.072
0.252
0.334
0.082
0.105
∙ CSDI ( 2021 )
0.029
0.050
0.328
0.742
-
-
-
-
-
-
∙ TSDiff ( 2023 )
0.036
0.057
0.324
0.731
0.057
0.070
0.243
0.330
0.079
0.091
∙ MG-TSD ( 2024 )
0.024
0.038
0.328
0.737
0.042
0.067
0.217
0.327
0.071
0.096
Table 1: The detailed results of ProtoFlow on the 48 and 96 prediction horizons. ProtoFlow-G: the variant of ProtoFlow that uses a standard Gaussian prior. - marks out-of-memory failures, Best and second-best forecasting results are bolded and underlined, respectively.
Method
Electricity
Solar
Traffic
Taxi
Wikipedia
Avg.
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
VAE + AR
0.421
0.544
1.154
2.618
1.005
1.511
0.337
0.448
0.130
0.169
0.609
1.058
VAE + FM
0.129
0.213
0.643
1.324
0.323
0.564
0.227
0.293
0.117
0.128
0.288
0.504
VQ + AR
0.037
0.054
0.347
0.833
0.066
0.084
0.205
0.322
0.108
0.121
0.153
0.283
ProtoFlow-G
0.019
0.037
0.305
0.717
0.036
0.058
0.159
0.287
0.096
0.114
0.123
0.243
ProtoFlow
0.015
0.022
0.284
0.638
0.034
0.054
0.157
0.281
0.092
0.101
0.116
0.219
Table 2: The comparison results of 96 prediction horizons with baselines regarding CRPSsum (C-S) and NRMSEsum (NM-S). AR: Auto-regressive Transformer and FM: Flow Matching .
Figure 2: Comparison of training convergence ( a ), inference efficiency, training memory ( b ), and forecasting performance ( c-d ) on Electricity, Traffic, and Solar with a prediction length of 96.
Method
Electricity
Solar
Traffic
Taxi
Wikipedia
Avg.
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
C-S
NM-S
ProtoFlow-G
0.019
0.037
0.305
0.717
0.036
0.058
0.159
0.287
0.096
0.114
0.123
0.243
ProtoFlow-uniform
0.018
0.030
0.356
0.769
0.036
0.056
0.157
0.284
0.099
0.108
0.133
0.249
ProtoFlow
0.015
0.022
0.284
0.638
0.034
0.054
0.157
0.281
0.092
0.101
0.116
0.219
Table 3: The forecasting performance with frequency-weighted prototype (ProtoFlow), uniform-weight prototype (ProtoFlow-uniform), and isotropic Gaussian priors (ProtoFlow-G) at a prediction length of 96.
Figure 3: Effects of codebook size on ProtoFlow at the 96-step prediction length.
Electricity
Solar
Traffic
Frequency
Uniform
Stage1 MSE
Frequency
Uniform
Stage1 MSE
Frequency
Uniform
Stage1 MSE
Coodebook size
64
0.039
0.044
7.76e-4
1.035
0.962
7.70e-4
0.113
0.105
5.82e-4
128
0.030
0.032
7.13e-4
0.774
0.761
6.75e-4
0.062
0.066
5.23e-4
256
0.022
0.030
6.02e-4
0.638
0.769
6.13e-4
0.054
0.056
4.78e-4
512
0.031
0.033
6.01e-4
0.993
0.758
6.18e-4
0.056
0.057
4.69e-4
Table 4: Effects of codebook size on Stage 1 reconstruction MSE and NRMSEsum under frequency and uniform weighting at a 96-step prediction length.
Figure 4: Effect of ODE sampling steps on 96-step probabilistic forecasting performance and W2 transport distance for ProtoFlow and ProtoFlow-G.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
DATASET
Dimension
Domain
Freq
Total Time Steps
Pred Length
Rolling Windows
SOLAR
137
R+
Hourly
7,009
{48,96}
7
ELECTRICITY
370
R+
Hourly
5,790
{48,96}
7
TRAFFIC
963
(0,1)
Hourly
10,413
{48,96}
7
TAXI
1214
N
30-Min
1,488
{48,96}
57
WIKIPEDIA
2000
N
Daily
792
{48,96}
5
Appendix
Table 5: The dataset properties used in our experiments are summarized as follows.
Datasets
Codebook Size
Codebook dim
Enc/Dec Layers
λ
β
Solar
256
128
2
0.99
0.1
Electricity
256
256
2
0.99
0.1
Traffic
256
256
2
0.99
0.1
Taxi
256
512
2
0.99
0.05
Wikipedia
256
512
2
0.99
0.05
Appendix
Table 6: Detailed hyperparameters of Stage 1.
Datasets
Hidden dim
DiTθ layers
Batch size
Steps
lr
Solar
128
4
64
3
5e-4
Electricity
256
4
64
3
5e-4
Traffic
256
5
64
3
3e-4
Taxi
512
5
64
3
3e-4
Wikipedia
512
4
64
5
3e-4
Appendix
Table 7: Detailed hyperparameters of Stage 2.
Layer
Function
Descriptions
1
Convolution
input channel= d , output channel=D, kernel size=3, stride=1, padding=1
Figure 6: ProtoFlow forecasts with a prediction horizon of 96. Rows show Electricity (370 dimensions), Traffic (963 dimensions), Solar (137 dimensions), Taxi (1214 dimensions), and Wikipedia (2000 dimensions) from top to bottom, with four examples per dataset.