Recent work has shown that probabilistic flow matching for time series forecasting benefits from a data-matched prior. The resulting prior introduces local correlations, which a sequential architecture usually absorbs: a recurrent neural network (RNN), a structured state-space model (S4), or a Transformer. However, such a backbone costs GPU memory and time per epoch. A cheaper alternative is MLP-based latent-space flow matching: embed the time series via an invertible map to a single latent vector and learn the flow there, so a tabular MLP can treat the series as a set of features. The relationship between the prior and the choice of linear latent map is understudied in conditional flow matching (CFM) forecasting, yet we found it strongly affects performance. Fixed transforms such as Fourier or discrete cosine (DCT) are only well-conditioned for Ornstein--Uhlenbeck priors, while a principal-component (PCA) map fit to the data is a strong but training-set-dependent reference sensitive to train--test shift. Instead, we propose to use the Mercer eigenbasis of the prior kernel: it diagonalises the centred covariance exactly, decouples from training data, and adapts to non-stationary and periodic priors. On five GluonTS benchmarks (ETTh1, ETTh2, Weather, Electricity, Traffic) under a shared protocol with TSFlow, the resulting MLP matches or beats it on CRPS at about 4.7× less training memory and 3.5×--4.4× less time per epoch.
Figures & tables
Figure 1: Overview of MercerFlow . The joint context and horizon window is mapped into the decorrelated Mercer eigenbasis of the GP prior. A lightweight MLP learns conditional flow matching directly in this latent space, enabling fast and scalable probabilistic forecasting without sequential backbones.
Ours
Linear Codec Baselines
Sequential
Metric
Mercer
Fourier
DCT
PCA
Raw
TSFlow (S4)
Mean CRPS rank ( ↓ )
1.00
1.20
1.40
1.80
3.40
5.80
# CRPS wins ( ↑ )
5/5
4/5
4/5
3/5
2/5
0/5
Mean MAE rank ( ↓ )
1.00
1.60
1.40
2.00
3.20
5.00
# MAE wins ( ↑ )
5/5
3/5
4/5
3/5
2/5
1/5
Table 1: Predictive performance summary across all five benchmarks (ETTh1, ETTh2, Weather, Electricity, Traffic). Rank =1+ number of methods strictly outperforming the model under a paired two-sided t -test on shared seeds ( p<0.2 , lower score is better); a win denotes rank 1 .
Metric / Architecture
ETTh1
ETTh2
Weather
Electricity
Traffic
MLP Time (s, ↓ )
0.67
0.69
1.80
0.93
0.68
S4 Time (s, ↓ )
2.50
2.40
7.70
4.10
2.60
MLP VRAM (GB, ↓ )
1.5
1.5
1.5
1.5
1.5
S4 VRAM (GB, ↓ )
7.1
7.1
7.1
7.1
7.1
Table 2: Computational efficiency benchmark at batch size 512. Comparison of per-epoch training time (seconds, ↓ ) and peak GPU memory (GB, ↓ ).
Figure 2: Traffic decorrelation on conditional-prior draws c0=Φ⊤[yp;y0] for the joint window. Each panel displays ∣DAD∣ for the leading 64 modes, with condition number κ(DAD) and test CRPS ( η=10−3 ) shown above each plot. Aligning the codec basis with the prior kernel maintains low off-diagonal correlation and preserves downstream performance.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Codec
ETTh1
ETTh2
Mercer (Ours)
0.2255±0.0007
0.145±0.002
Fourier
0.2240±0.0020
0.145±0.003
DCT
0.2260±0.0010
0.145±0.002
PCA
0.2248±0.0004
0.145±0.002
Raw
0.2240±0.0010
0.145±0.003
Appendix
Table 3: SGD ablation on ETTh1 and ETTh2 (Test CRPS, ↓ ) for Post-AdaLN Nesterov SGD (momentum 0.9 , learning rate 10−3 ). Evaluated at 20 NFE under the identical joint window ( L=384 ) and OU length scales as the AdamW runs.
Codec
First-layer Gram Matrix A
Jacobi Condition Number Bound κ(DAD)
Identity (Raw)
Tn(f)+εIn
κ(A) , increases with length scale ℓ
Mercer (Ours)
Λ=diag(λ1,…,λn)
1 (exact decorrelation)
Fourier
ΦF⊤Tn(f)ΦF+εIn
1+O(1/minωf(ω))
DCT
ΦD⊤Tn(f)ΦD+εIn
1+O(1/minωf(ω))
PCA
Empirical data covariance
Data-dependent (not aligned with prior kernel)
Appendix
Table 4: Theoretical first-layer Gram matrix A=E[xx⊤] and corresponding Jacobi condition number κ(DAD) under the centred Ornstein–Uhlenbeck prior y∼N(0,K) on a window of length n=L , with regularizing jitter ε>0 .
Dataset
Raw
Fourier
Mercer
DCT
PCA
ETTh1
1.72×105
268.50
1.00
1.92
8.15
ETTh2
8.16×104
189.71
1.00
2.07
46.98
Weather
2.80×104
119.74
1.00
2.12
13.46
Electricity
2.23×103
41.63
1.00
2.08
43.93
Traffic
195.79
13.44
1.00
1.90
44.68
Appendix
Table 5: Empirical Jacobi condition number κ(DAD) under the centred prior y∼N(0,K) on the joint grid (joint window 192 / 192 ; last channel; scaler fit on the first 70% ). This is the population the proposition’s identity is proved against: Mercer hits κ(DAD)=1 on every dataset, with jitter only, by construction. The training input is not this population; Table 12 reports the same quantity under the conditional draw c0=Φ⊤[yp;y0] .
Dataset
OU
RBF
Periodic
ETTh1
−1.231
−0.970
0.062
ETTh2
−0.836
−1.184
0.608
Traffic
0.465
0.226
0.035
Electricity
0.228
0.209
0.969
Weather
−0.587
−0.372
1.203
Appendix
Table 6: Mean negative log density of candidate Gaussian process priors on the validation slice. Lower is better ( ↓ ). The periodic entry for Traffic uses period p=168 , length scale ℓ=0.1 , and jitter ε=0.05 ; all other entries represent the optimal fit across 400 TPE trials. Bold indicates the best score in each row.
Dataset
Kernel
ℓ
Period p
Jitter ε
ETTh1
OU
479
—
4.4×10−4
ETTh1
RBF
7.41
—
2.8×10−3
ETTh1
Periodic
1.98
276
0.062
ETTh2
OU
183
—
1.0×10−5
ETTh2
RBF
4.58
—
7.2×10−4
ETTh2
Periodic
2.77
24
0.23
Appendix
Table 7: Continuous optimal parameters of the validation likelihood fits from Table 6 . Length scale ℓ , period p (periodic kernel only), and regularizing diagonal jitter ε . The jitter search interval is ε∈[10−5,1.0] ; the OU fits for ETTh2, Traffic, Electricity, and Weather saturate at the lower bound 1.0×10−5 .
Dataset
Length Scale ℓ
Runtime Jitter ε
ETTh1
336
10−4
ETTh2
192
10−4
Weather
96
10−4
Electricity
24
10−4
Traffic
7
10−4
Appendix
Table 8: Baseline OU length scale ℓ and runtime diagonal jitter ε used in primary benchmark runs, shared across Mercer, Fourier, DCT, PCA, and Raw.
Prior
Mercer
Fourier
Raw
DCT
PCA
OU
687
891
8.54×103
712
477
Periodic
1.18×103
188
1.62×104
2.45×104
23.0
Appendix
Table 9: Traffic first-layer Jacobi condition number κ(DAD) under conditional prior samples c0=Φ⊤[yp,y0] : stationary OU prior ( ℓ=7 , ε=10−4 ) versus validation-optimal periodic prior (period p=168 , ℓ=0.1 , ε=0.05 ). Evaluated on the joint window ( Lp=192,Lf=192 , L=384 ), last channel, random seed 42 , standard scaler fit on the first 70% . Bold indicates the best (lowest) condition number in each row.
Prior
Learning Rate η
Mercer
Fourier
Raw
DCT
PCA
OU
10−3
0.0875±0.0009
0.0871±0.0015
0.0884±0.0003
0.0878±0.0011
0.0883±0.0018
OU
10−1
0.1065±0.0036
0.1080±0.0086
0.2462±0.0717
0.1116±0.0026
0.1065±0.0041
Periodic
10−3
0.0821±0.0006
0.0827±0.0011
0.0839±0.0016
0.0840±0.0007
0.0828±0.0007
Periodic
10−1
0.0951±0.0130
0.7773±1.0350
1.0198±0.1078
0.1641±0.0688
0.0855±0.0032
Appendix
Table 10: Traffic probabilistic forecasting performance (Test CRPS, 20 NFE, ↓ ) under the stationary OU prior vs. the validation-optimal periodic prior ( p=168 , ℓ=0.1 , ε=0.05 ). Under the periodic prior, MercerFlow dynamically adapts and achieves top performance, whereas fixed harmonic codecs (DCT, Fourier) suffer from severe condition degradation and numerical instability under learning rate stress ( η=10−1 ). Bold indicates the best score in each row.
ETTh2
Weather
Electricity
ETTh1
Traffic
OU NLL
−0.836
−0.587
0.228
−1.231
0.465
Residual
0.33
0.44
1.00
1.44
1.46
Appendix
Table 11: Dataset hardness from the frozen OU prior. Residual is the mean squared error of the conditional OU mean on the held-out last 15% , after the training standard scaler, averaged over channels; length scale from Table 8 . OU NLL is the validation score from Table 6 ; white noise on unit-scaled series scores about 1.42 . Easy: residual <1/2 ; hard: residual ≥1 .
Dataset
Raw
Fourier
Mercer
DCT
PCA
ETTh1
(2.08/3.25) 1e5
539 / 471
34.7 / 14.8
34.7 / 14.6
26.8 / 15.1
ETTh2
(3.45/6.21) 1e5
1770 / 2970
155 / 248
154 / 260
121 / 198
Weather
(4.20/1.11) 1e4
295 / 272
13.5 / 4.5
14.3 / 4.6
17.6 / 15.9
Electricity
(5.85/5.22) 1e3
234 / 394
60.2 / 171
59.6 / 176
115 / 285
Traffic
(8.60/1.10) 1e3
922 / 1140
683 / 925
709 / 936
481 / 1290
Appendix
Table 12: Empirical Jacobi condition number κ(DAD) under the conditional prior draws c0=Φ⊤[yp;y0] on the training split and on the held-out last 15% , standard-scaled; D=diag(A)−1/2 with A the empirical second moment of c0 . Each cell gives train/test. Scaler fit on the first 70% ; the PCA basis is fit on data windows, not on prior draws. Raw values share a single exponent per row.
κ(DAD)
Overall ( n=5 )
Hard ( n=3 )
Easy ( n=2 )
Test
0.44
0.79
−0.08
Train
0.38
0.68
−0.08
Appendix
Table 13: Mean across datasets of within-dataset Spearman ρ between κ(DAD) under the conditional prior (Table 12 ) and CRPS, over the five MLP codecs; TSFlow excluded. Each per-dataset ρ ranks the five codecs on the two axes within that dataset; the column is the mean of those ρ ’s across the listed datasets. n is the number of datasets averaged.
Method
ETTh1
ETTh2
Weather
Electricity
Traffic
DLinear ( Zeng et al., 2023 )
23.06±0.01
16.02±0.01
6.26±0.02
5.57±0.00
10.75±0.01
PatchTST ( Nie et al., 2022 )
23.02±0.61
16.50±0.47
5.99±0.58
6.70±0.20
12.45±0.30
DMamba ( Chen and Sun, 2026 )
25.38±0.89
16.00±0.41
2.56±0.02
5.48±0.10
10.42±0.24
TSFlow (S4) ( Kollovieh et al., 2025 )
19.1±0.2 (5)
15.4±0.3 (6)
2.18±0.08 (6)
7.4±0.3 (6)
23.0±9.0 (6)
Raw
19.2±0.2 (5)
13.1 ± 0.2 (1)
1.82±0.03 (1)
4.23±0.06 (5)
8.9±0.1 (5)
Fourier
19.0±0.3 (1)
13.2±0.3 (1)
1.82±0.03 (1)
4.11±0.04 (2)
8.70 ± 0.09 (1)
Appendix
Table 14: Comprehensive long-horizon probabilistic forecasting performance (Test CRPS, 10−2 , ↓ ). Top: deterministic baselines; bottom: conditional flow matching models. Parenthetical values denote paired two-sided t -test rank ( p<0.2 ) among flow matching models on shared random seeds ( 42 – 48 ). Bold : best; second best .
Method
ETTh1
ETTh2
Weather
Electricity
Traffic
DLinear ( Zeng et al., 2023 )
2.12±0.01
4.91±0.01
27.56±0.08
187.2±0.2
4.00±0.01
PatchTST ( Nie et al., 2022 )
2.11 ± 0.06
5.06±0.15
26.38±2.54
225.4±6.8
5.00±0.01
DMamba ( Chen and Sun, 2026 )
2.33±0.08
4.91 ± 0.12
11.26±0.10
184.2±3.5
4.00±0.01
TSFlow (S4) ( Kollovieh et al., 2025 )
2.25±0.03 (1)
6.0±0.1 (6)
11.6±0.3 (6)
320±10 (6)
11.0±5.0 (6)
Raw
2.31±0.02 (6)
5.13±0.09 (1)
10.0 ± 0.2 (1)
181±3 (5)
4.04±0.08 (3)
Fourier
2.27±0.03 (3)
5.1±0.1 (1)
10.1±0.2 (1)
176±2 (2)
3.95 ± 0.04 (1)
Appendix
Table 15: Point forecasting performance (Test MAE, ↓ , Traffic in 10−3 ). Parenthetical values denote paired two-sided t -test rank ( p<0.2 ) among flow matching models on shared random seeds ( 42 – 48 ). Bold : best; second best .
Lf
Method
ETTh1
ETTh2
Weather
Electricity
Traffic †
CRPS ( ×10−2 , ↓ )
336
Raw
17.77 ± 0.23
13.07 ± 0.33
3.98 ± 0.05
4.43 ± 0.09
10.36 ± 0.20
Fourier
18.07 ± 0.49
12.91 ± 0.19
3.82 ± 0.02
4.51 ± 0.10
10.41 ± 0.15
DCT
17.81 ± 0.67
12.65 ± 0.31
3.82 ± 0.02
4.49 ± 0.09
10.45 ± 0.15
PCA
17.93 ± 0.49
12.51 ± 0.30
3.83 ± 0.03
4.49 ± 0.08
10.29 ± 0.16
MercerFlow
17.82 ± 0.49
12.66 ± 0.36
3.82 ± 0.03
4.52 ± 0.10
10.41 ± 0.21
Appendix
Table 16: Sensitivity to extended forecasting horizons ( Lp=192 ). Bold: best; underline: second best.
Dataset
Mercer (Ours)
Fourier
DCT
PCA
Raw
ETTh1
17.99±0.44
18.17±0.55
18.51±0.33
17.81±0.00
122.42±80.10
ETTh2
13.05±0.56
13.27±0.06
14.00±1.13
12.92±0.20
15.18±0.69
Weather
1.70±0.01
1.71±0.01
1.71±0.01
1.70±0.01
1.71±0.06
Electricity
4.34±0.06
4.33±0.06
4.34±0.04
4.39±0.05
6.53±0.35
Traffic
10.65±0.36
10.80±0.86
11.16±0.26
10.65±0.41
24.62±7.17
Appendix
Table 17: Test CRPS ( ×10−2 , ↓ ) at AdamW learning rate η=10−1 (default is 10−3 ), mean ± std over seeds 42 – 44 . Evaluated at matched mid-run EMA checkpoints E∗ . Bold: within one std of the best.
Figure 3: Training loss trajectories on ETTh1 and ETTh2 at an elevated learning rate ( η=10−2 ). MercerFlow (solid blue) converges faster and strictly monotonically. In contrast, the unrotated baseline ( RawFlow , dashed orange) suffers from optimization friction and visible instability (e.g., the loss spike around epoch 60 on ETTh1), demonstrating the empirical benefits of first-layer preconditioning.
Codec / arm
κ(DAD) prior
κ(DAD) interp.
Test CRPS ( ↓ )
SGD gap (%)
Mercer ( ℓ=32 )
1.00
98.5
0.6063±0.0025
0.00
Fourier ( ℓ=32 )
24.90
654.0
0.6064±0.0023
0.00
DCT ( ℓ=32 )
1.82
105.0
0.6062±0.0024
0.00
Raw ( ℓ=32 )
1,501.0
3,515.0
0.6069±0.0030
0.00
Mercer ( ℓ=320 )
1.71
192.0
0.6125±0.0057
—
Appendix
Table 18: Synthetic OU comparison (context/horizon 32/32 , data length scale ℓ=32 ). Matched codecs share the OU prior at ℓ=32 ; the wrong- ℓ arm uses a Mercer/GP model at ℓ=320 . AdaLN MLP ( 64 hidden, 2 layers), AdamW η=10−3 , 2,000 steps, seeds 0 – 2 (mean ± std). κ(DAD) prior is computed on the population covariance Φ⊤KΦ ; interpolant κ(DAD) is empirical on conditional CFM states ct . SGD control: Nesterov momentum 0.9 , matched compute budget; gaps are 100×(CRPS/CRPSRaw−1) (mean over seeds). Bold indicates values within one standard deviation of the best.
NFE (Euler Steps)
Dataset
Method
4
10
20
50
128
ETTh1
MercerDiff
19.03±0.07
1.04±0.06
0.56±0.07
0.51±0.06
0.47±0.03
MercerFlow
0.200±0.017
0.199±0.017
0.199±0.017
—
—
ETTh2
MercerDiff
7.78±0.02
0.47±0.01
0.211±0.005
0.176±0.005
0.177±0.003
MercerFlow
0.137±0.005
0.136±0.006
0.136±0.007
—
—
Appendix
Table 19: Inference sampling efficiency: Test CRPS ( ↓ , mean ± std) across varying Numbers of Function Evaluations (NFE) at horizon Lf=192 . Due to straight probability paths, MercerFlow saturates within 4 – 10 steps, whereas diffusion ( MercerDiff ) suffers from severe discretization error. Bold: best per column.
Diagonal Jitter ε
Metric
Dataset
Codec
10−4
10−3
10−2
10−1
1.0
κ(DAD) ( ct , ↓ )
ETTh1
Mercer
23.55
30.64
29.37
47.03
334.67
Raw
2.01×105
3.02×105
1.85×105
1.66×105
1.44×105
ETTh2
Mercer
62.14
74.75
95.27
524.26
3.64×103
Raw
4.07×105
1.22×106
1.20×106
1.17×106
1.01×106
CRPS ( ↓ )
ETTh1
Mercer
0.188
0.186
0.187
0.195
0.255
Appendix
Table 20: Sensitivity to GP prior diagonal jitter ε on ETTh1 and ETTh2 ( L=384 ). κ(DAD) is evaluated on conditional interpolants ct . Standard regularisation ( ε≤10−2 ) preserves well-conditioned representations; extreme jitter ( ε≥0.1 ) collapses the prior toward white noise ( K≈IL ), degrading predictive accuracy. Bold: best per column within each dataset block.
Department of Computer Science and Engineering, University of Science and Technology Chittagong · Department of Electrical and Electronic Engineering, University of Science and Technology Chittagong · Faculty of Science, Engineering and Technology, University of Science and Technology Chittagong