Many scientific forecasting tasks require updating a distribution over future trajectories as new observations arrive. Conventional diffusion and flow models generate each forecast from Gaussian noise, often at the cost of many sampling steps. Warm-start methods reuse earlier predictions to reduce this cost, but their models are not trained to perform the forecast update itself, which can compromise quality under few-step sampling. In this work, we introduce Seq-Flow, a conditional flow model whose ODE transports samples from the previous forecast distribution to the updated one. Because successive forecasts often differ only modestly, this transport starts from an informative distribution and can produce accurate updates with few flow evaluations. Recursive reuse also creates a challenge: errors in one forecast become errors in the initial states of subsequent flows. We address this with self-rollout training, in which a moving average copy of the model generates forecasts that initialize later training updates. Unlike self-forcing methods, which reuse generated outputs as conditioning context, Seq-Flow reuses them as the source of the next flow. Experiments On particle-accelerator beam spill forecasting show Seq-Flow reduces CRPS by 65% under a few-NFE sampling budget, while remaining competitive with strong baselines on fluid-dynamics forecasting tasks. Although trained on self-rollouts of at most four updates, Seq-Flow remains accurate over more than 400 consecutive updates. Our code is available at https://github.com/Graph-COM/Seq-Flow.
Figures & tables
Figure 1: Seq-Flow with self-rollout training. Seq-Flow tracks the forecast distribution by modeling the forecast update. It uses a renoised sample from the previous forecast as the source distribution for the next flow, enabling efficient few-step updates. Self-rollout training exposes the model to its own forecast sources, teaching it to correct errors that would otherwise compound across updates.
Method
NFE
Synthetic
Beam Spill
Burgers’ Equation
RealPDE-FSI
W1 dist. ↓
CRPS ↓
RMSE ↓
CRPS ↓
RMSE ↓
CRPS ↓
RMSE ↓
AR Diffusion
H×3
1.270
0.363
1.001
0.0589
0.1183
0.0150
0.0260
Self-forced AR Diffusion
H×3
2.765
0.055
0.209
0.0143
0.0395
0.0031
0.0062
Full-trajectory Diffusion
3
2.260
0.043
0.179
0.0139
0.0366
0.0030
0.0063
AR Flow
H×3
1.203
0.446
1.092
0.0822
0.1401
0.0130
0.0208
Self-forced AR Flow
H×3
2.872
0.059
0.230
0.0198
0.0356
0.0039
0.0081
Table 1: Results on the synthetic dataset, Beam Spill, Burgers’ Equation, and RealPDE-FSI. Prediction horizon H=5 for the synthetic task and 10 for the remaining tasks.
Figure 2: Performance-efficiency trade-offs . X axis: sampling steps (NFE), Y axis: task performance. For clarity, each figure shows only a subset of the top-performing baselines. Left: synthetic task. Right: beam spill forecasting.
Figure 3: Seq-Flow with controlled error accumulation over forecast steps . Each figure shows performance over forecast step (window id). Left: synthetic task. Right: beam spill forecasting.
Figure 4: Short self-rollout steps T0 leads to robust long-run forecasting . Each figure shows performance per forecast step with different T0 . Left: synthetic task. Right: beam spill forecasting.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
State shape
Training
Held-out
L
H
Synthetic
1
50,000
5,000
100
5
Beam Spill
1
40,000
5,000
430
10
Burgers
32
90,000
10,000
100
10
RealPDE-FSI
2×32×32
39
6
100
10
Appendix
Table 2: Dataset settings. L : episode length; H : forecast horizon. Counts are trajectories; field shapes are channels × grid.
Transformer
Spatiotemporal U-Net
Synthetic, Beam Spill, Burgers
RealPDE-FSI
Hidden width
128
Base width
64
Layers
12
Channel multipliers
(1,2,4)
Feed-forward width
512
Attention heads
4
Attention heads
4
Head dimension
32
Appendix
Table 3: Backbone configurations. U-Net channel multipliers scale the base width.
Probabilistic forecasting is important for predicting complex dynamical systems because intrinsic randomness and incomplete observations can cause the same observed state to evolve into multiple plausible futures. While flow matching is a flexible approach for probabilistic forecasting, it is computationally expensive. Streaming flow (SF) reformulates this approach to model temporal evolution efficiently by learning a continuous velocity field directly in physical time. However, SF learns a deterministic velocity field. Thus, it provides only a single future trajectory for a given fixed initial state and observation history. To overcome this limitation, we introduce Variational Streaming Flow (VSF). Our approach learns a latent distribution that is conditioned on the dynamics of interest. In turn, this enables probabilistic forecasting. Importantly, we retain the computational efficiency of SF by generating in physical time. Across deterministic and stochastic dynamical systems, VSF demonstrates superior predictive accuracy and distributional fidelity. We demonstrate the advantage for both long-horizon rollouts exceeding 1,000 steps, and settings with bifurcating dynamics. Moreover, VSF can be integrated into existing Joint-Embedding Predictive Architecture (JEPA)-based world models as a plug-and-play predictor to improve temporal dynamics and goal-directed success rate in navigation, motion planning, and manipulation.
Hans Hao-Hsun Hsu, Minseon Gwak, Soon Hoe Lim +2
Georgia Institute of Technology · Lawrence Berkeley National Lab · International Computer Science Institute +2
Deep learning surrogates for CFD flow-field prediction often rely on large, complex models, which can be slow and fragile when data are noisy or incomplete. We introduce FlowForge, a staged local rollout engine that predicts future flow fields by compiling a locality-preserving update schedule and executing it with a shared lightweight local predictor. Rather than producing the next frame in a single global pass, FlowForge rewrites spatial sites stage by stage so that each update conditions only on bounded local context exposed by earlier stages. This compile-execute design aligns inference with short-range physical dependence, keeps latency predictable, and limits error amplification from global mixing. Across PDEBench, CFDBench, and BubbleML, FlowForge matches or improves upon strong baselines in pointwise accuracy, delivers consistently better robustness to noise and missing observations, and maintains stable multi-step rollout behavior while reducing per-step latency.
Xiaowen Zhang, Ziming Zhou, Fengnian Zhao +1
Shanghai Jiao Tong University · University of Michigan
Fast surrogate modeling for high-dimensional physical dynamics requires more than low short-term error: useful models must roll out efficiently while preserving the statistical structure of long trajectories. Neural operators provide inexpensive autoregressive forecasts but can drift in turbulent regimes, whereas rolling diffusion and latent generative surrogates can represent stochastic transitions at the cost of multi-step denoising, noise-schedule design, or auxiliary compression models. We propose MeanFlow Long-term Invariant Spatiotemporal Consistency Autoregressive Models (MeLISA), a latent-free autoregressive generative surrogate built on pixel-space MeanFlow. MeLISA defines a blockwise stochastic transition kernel that generates each forecast block with a single model evaluation, avoiding latent encoders and iterative diffusion solvers at inference time. To stabilize long-horizon rollouts, MeLISA combines a Window-Consistency MeanFlow objective that learns conditional spatiotemporal generation from partially observed temporal windows with a Time Increment Consistency loss that constrains multi-lag finite increments and targets temporal-correlation structure. We evaluate MeLISA with compact UNet and scalable DiT backbones on two high-resolution benchmarks, extended 2D Kolmogorov flow at 256×256 and turbulent channel-flow slice at 192×192. MeLISA outperforms neural-operator baselines on short-term forecasting accuracy and long-horizon statistical metrics, including energy spectra, turbulent kinetic energy, and mixing-rate-related dynamics, while achieving inference speeds comparable to, and in some cases faster than, neural operators. Compact 3.7-5.7M-parameter variants already deliver strong parameter efficiency, and DiT variants provide a scalable path up to 150M parameters. Overall, MeLISA benefits both rollout efficiency and long-horizon statistical accuracy.
Tianyue Yang, Xiao Xue
The Center for Computational Science University College London