We introduce CNP-Flow, a flow matching framework for temporal generation that learns conditional source distributions through flow reversal. Whereas standard conditional flow matching (FM) incorporates conditioning through the vector field and draws source samples from a standard Gaussian, CNP-Flow uses a conditional noise predictor (CNP) to produce an isotropic Gaussian source for each temporal condition. The CNP is supervised by source samples obtained through flow reversal, which maps observed targets backward through a pretrained FM model. A three-stage pipeline pretrains the FM model, trains the CNP, and fine-tunes the FM model using the learned source distribution, while preserving the FM backbone architecture. Across video prediction, video interpolation, and 7-DoF Franka robot motion planning, CNP-Flow consistently improves generation quality. It also matches baseline performance with fewer function evaluations. Project page: https://embodiedai-ntu.github.io/cnpflow
Figures & tables
Figure 1: CNP-Flow learns condition-dependent source distributions for temporal generation. CNP-Flow uses a Conditional Noise Predictor (CNP) to predict a source distribution for each condition. This source-side conditioning reduces FVD by up to 75% over RIVER on video prediction, and yields more precise motion-planning trajectories with fewer collisions.
Figure 2: 2D toy example of temporal generation . Colors indicate temporal phase along the target circular trajectory. (a) Vanilla FM draws source samples from a standard Gaussian. Samples associated with neighboring temporal conditions can be widely separated, and the generated sequence deviates from the target trajectory. (b) CNP-Flow learns a conditional source distribution with a smoother arrangement of samples by temporal phase (see the zoomed region), and the generated sequence follows the target trajectory more closely. Detailed experimental settings are provided in Appendix F .
Figure 3: Overview of the proposed three-stage training pipeline. Stage 1: A pretrained FM is used to obtain flow-reversal-induced source samples, defining a source distribution aligned with the target data. Stage 2: The CNP is trained to model this distribution by predicting a conditional Gaussian prior via NLL and KL objectives. Stage 3: The FM model is fine-tuned using source samples drawn from the learned conditional prior, resulting in more coherent and target-aligned transport trajectories.
Dataset
p→k
Method
PSNR ↑
SSIM ↑
FVD ↓
BAIR64
1→15
RIVER ( Davtyan et al., 2023 )
–
–
73.50
MAGVIT ( Yu et al., 2023 )
19.30
0.787
62.00
iVideoGPT ( Wu et al., 2024 )
20.40
0.823
75.00
CAR-Flow ( Chen et al., 2025 )
20.08
0.836
71.37
PYoCo ( Ge et al., 2023 )
20.13
0.837
61.88
CSFM ( Kim et al., 2026 )
20.16
0.836
64.20
Table 1: Results on video prediction. CNP-Flow improves over the fixed-source RIVER baseline and remains competitive with recent source-learning methods. CNP-Flow results are averaged over three seeds.
Method
Success ↑
Valid (%) ↑
EE pos (cm) ↓
EE ori ( ∘ ) ↓
Collision ↓
MPD ( Carvalho et al., 2025 )
97%
86.6
1.9
1.2
0.0030
Fixed-source FM
96%
84.9
1.6
1.1
0.0038
CNP-Flow (Ours)
97%
88.5
0.8
0.5
0.0023
Table 3: Results on motion planning. All methods use a matched 15-step sampling budget, and results are averaged over three seeds.
Figure 4: Inference efficiency on KTH prediction. Curves show PSNR/SSIM averaged across generated samples per video. CNP-Flow reaches the standard Gaussian PSNR plateau ( ≈26.5 ) at NFE =5 and converges higher ( ≈28.2 ) by NFE =10 ; SSIM shows a similar trend. The crossover point is highlighted on each panel.
Figure 5: Qualitative comparison. Left : CLEVRER interpolation. Fixed-source FM drifts at the highlighted frames, while CNP-Flow matches the ground truth. Right : Franka Panda motion planning. Each panel overlays 50 sampled trajectories (red: colliding, green: collision-free). CNP-Flow produces more collision-free plans than MPD and fixed-source FM.
Video prediction (KTH)
Source x0
PSNR ↑
SSIM ↑
FVD ↓
N(0,I)
30.40
0.860
180.00
μϕ(Ci)
31.46
0.908
83.82
N(μϕ(Ci),σϕ2(Ci)I)
34.25
0.943
44.76
Table 4: Ablation studies. Left : source parameterization on KTH 10→30 video prediction, isolating the contribution of the learned mean and learned variance. Right : source learning and backbone fine-tuning (FT) on Franka Panda motion planning. In both tasks, metrics improve as the source better matches the condition.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Fixed-source FM
CNP-Flow
Latency increase
Additional memory
KTH / BAIR64
207.1 ms/frame
214.2 ms/frame
3.43%
∼15 MB
CLEVRER
211.9 ms/frame
219.1 ms/frame
3.39%
Appendix
Table 5: Inference overhead of CNP-Flow. Latency is measured over 1,000 generated frames under matched hardware, batch-size, and solver settings.
Component
KTH
CLEVRER
BAIR64
Flow reversal steps
100
20
50
FM fine-tune length
200k steps
20k steps
50k steps
FM fine-tune batch
64
128
64
FM fine-tune LR
5×10−5
5×10−5
5×10−5
Appendix
Table 6: Video prediction fixed-source FM and flow reversal configuration. Source samples obtained through flow reversal are cached only for CNP supervision and are not used at inference time.
Component
KTH
CLEVRER
FM pretrain steps
300k
300k
FM pretrain batch
32
16
Flow reversal steps
100
40
FM fine-tune steps
200k
200k
FM fine-tune batch
32
16
FM fine-tune LR
1×10−4
1×10−4
Appendix
Table 7: Video interpolation fixed-source FM and flow reversal configuration. Since no compatible interpolation checkpoint is available, the fixed-source CFM backbone is trained from scratch.
Component
Setting
Backbone
Convolutional encoder
Input
Conditional latents + temporal-distance embedding
Output
μ,logσ2
Base channels
64
Channel multipliers
(1,2,4,8)
Residual blocks per resolution
2
Appendix
Table 8: Video CNP architecture. The same architecture is used for video prediction and video interpolation.
Task
Dataset
Length
Batch
KL warm-up
Prediction
KTH
50 epochs
64
10 epochs
Prediction
CLEVRER
15 epochs
1024
10 epochs
Prediction
BAIR64
15 epochs
64
5 epochs
Interpolation
KTH
200k steps
32
5 epochs
Interpolation
CLEVRER
200k steps
16
2 epochs
Appendix
Table 9: Video CNP training configuration. All video CNPs are trained with Adam using learning rate 1×10−4 and KL weight λKL=1×10−4 .
Dataset
PSNR ↑
SSIM ↑
FVD ↓
Video prediction
BAIR64 1→15
20.15±0.13
0.8390±0.0021
61.17±1.58
KTH 10→30
34.25±0.26
0.9432±0.0025
44.76±1.53
KTH 10→40
32.69±0.22
0.9303±0.0040
54.88±3.85
CLEVRER 2→14
42.23±0.19
0.9935±0.0002
24.18±2.46
Video interpolation
Appendix
Table 10: Variability across runs for CNP-Flow. Results are reported as mean ± standard deviation over three independent runs for video prediction and interpolation.
Component
Setting
Backbone
Residual MLP
Input
Start state + EE goal pose + context embedding
Raw condition dimension
26
Context embedding dimension
128
Output
μ,logσ2
Hidden dimension
1024
Appendix
Table 11: Motion planning CNP architecture. The CNP predicts an isotropic Gaussian source distribution over the learnable B-spline control-point noise.
Method
Valid (%) ↑
EE pos. (cm) ↓
EE ori. ( ∘ ) ↓
Collision ↓
Diffusion prior
64.3±1.5
5.3±0.1
7.5±0.5
0.0402±0.0013
Fixed-source FM
61.1±3.8
3.8±0.2
5.6±0.5
0.0445±0.0034
CNP-Flow
64.3±4.6
2.1±0.2
2.9±0.3
0.0366±0.0045
Appendix
Table 12: Motion-planning performance without post-hoc refinement. Results are reported as mean ± standard deviation over three independent runs.
Regime
Task
Split
Method
MSE ↓
M-MSE ↓
PSNR ↑
SSIM ↑
Rigid
push_cube
IND
Vanilla FM
0.002886
0.017354
26.29
0.9546
CNP-Flow (Ours)
0.002192
0.013373
27.80
0.9647
OOD
Vanilla FM
0.002977
0.016039
26.78
0.9539
CNP-Flow (Ours)
0.002338
0.012326
28.13
0.9629
Deformable
push_rope
IND
Vanilla FM
0.000213
0.002708
36.75
0.9878
CNP-Flow (Ours)
0.000207
0.002608
36.88
0.9889
Appendix
Table 13: Results on the ACWM-Phys physical world-model benchmark at 240×240 resolution. We evaluate one task from each physical regime on both in-distribution (IND) and out-of-distribution (OOD) splits. All results are averaged over three independent seeds.
Source, frozen FM
PSNR ↑
SSIM ↑
FVD ↓
Standard Gaussian
30.29
0.8591
179.97
CNP w/o flow reversal supervision
30.46
0.8612
179.54
CNP w/ flow reversal supervision
30.85
0.9008
100.91
Appendix
Table 14: Effect of flow reversal supervision on source learning. The Stage-1 FM is frozen for all variants.
Source
Pairwise LPIPS ↑
Standard Gaussian
0.0666
CNP-Flow
0.0686
Appendix
Table 15: Prediction diversity on BAIR64. We report the average pairwise LPIPS among predictions generated from the same condition.
Dataset
Round 1
Round 2
Round 3
BAIR64
61.17
60.48
67.07
KTH
44.76
35.04
36.37
Appendix
Table 16: Iterative Stage-2/Stage-3 refinement. FVD is reported after each round of source learning and FM refinement.
Flow reversal solver
FVD ↓
Dopri
61.73±1.47
Heun-100
61.16±0.99
Euler-100
61.17±1.56
Appendix
Table 17: Robustness to the numerical solver used for flow reversal on BAIR64 prediction. Values are mean ± standard deviation over three seeds.
Method
∥x0−x1∥2↓
Path length ↓
Turning (rad) ↓
Endpoint dev. ↓
Standard Gaussian
19.1715
18.6780
0.8292
2.7362
CNP-Flow (Ours)
17.4091
17.0088
0.7586
2.4710
Appendix
Table 18: Transport geometry on BAIR64. Both methods are evaluated on the same 256 condition–target pairs using Euler-100.
Coupling
Distance ↓
Path ↓
Turning ↓
Endpoint dev. ↓
Flow reversal pair
15.7342
15.7185
0.2261
0.4913
Conditional OT
17.3715
16.2187
1.0610
3.9173
CNP-Flow (Ours)
17.4091
17.0088
0.7586
2.4710
Appendix
Table 19: Alternative Stage-3 couplings on BAIR64 prediction. Left : transport geometry on 256 condition–target pairs using Euler-100. Right : downstream generation quality. Direct pairing via flow reversal requires the target and is therefore reported only as a training-time geometry reference.
Figure 6: Toy example visualization of the flow reversal mapping and transport distance analysis.