We introduce CNP-Flow, a flow matching framework for temporal generation that learns conditional source distributions through flow reversal. Whereas standard conditional flow matching (FM) incorporates conditioning through the vector field and draws source samples from a standard Gaussian, CNP-Flow uses a conditional noise predictor (CNP) to produce an isotropic Gaussian source for each temporal condition. The CNP is supervised by source samples obtained through flow reversal, which maps observed targets backward through a pretrained FM model. A three-stage pipeline pretrains the FM model, trains the CNP, and fine-tunes the FM model using the learned source distribution, while preserving the FM backbone architecture. Across video prediction, video interpolation, and 7-DoF Franka robot motion planning, CNP-Flow consistently improves generation quality. It also matches baseline performance with fewer function evaluations. Project page: https://embodiedai-ntu.github.io/cnpflow
Figures & tables
Figure 1: CNP-Flow learns condition-dependent source distributions for temporal generation. CNP-Flow uses a Conditional Noise Predictor (CNP) to predict a source distribution for each condition. This source-side conditioning reduces FVD by up to 75% over RIVER on video prediction, and yields more precise motion-planning trajectories with fewer collisions.
Figure 2: 2D toy example of temporal generation . Colors indicate temporal phase along the target circular trajectory. (a) Vanilla FM draws source samples from a standard Gaussian. Samples associated with neighboring temporal conditions can be widely separated, and the generated sequence deviates from the target trajectory. (b) CNP-Flow learns a conditional source distribution with a smoother arrangement of samples by temporal phase (see the zoomed region), and the generated sequence follows the target trajectory more closely. Detailed experimental settings are provided in Appendix F .
Figure 3: Overview of the proposed three-stage training pipeline. Stage 1: A pretrained FM is used to obtain flow-reversal-induced source samples, defining a source distribution aligned with the target data. Stage 2: The CNP is trained to model this distribution by predicting a conditional Gaussian prior via NLL and KL objectives. Stage 3: The FM model is fine-tuned using source samples drawn from the learned conditional prior, resulting in more coherent and target-aligned transport trajectories.
Dataset
p→k
Method
PSNR ↑
SSIM ↑
FVD ↓
BAIR64
1→15
RIVER ( Davtyan et al., 2023 )
–
–
73.50
MAGVIT ( Yu et al., 2023 )
19.30
0.787
62.00
iVideoGPT ( Wu et al., 2024 )
20.40
0.823
75.00
CAR-Flow ( Chen et al., 2025 )
20.08
0.836
71.37
PYoCo ( Ge et al., 2023 )
20.13
0.837
61.88
CSFM ( Kim et al., 2026 )
20.16
0.836
64.20
Table 1: Results on video prediction. CNP-Flow improves over the fixed-source RIVER baseline and remains competitive with recent source-learning methods. CNP-Flow results are averaged over three seeds.
Method
Success ↑
Valid (%) ↑
EE pos (cm) ↓
EE ori ( ∘ ) ↓
Collision ↓
MPD ( Carvalho et al., 2025 )
97%
86.6
1.9
1.2
0.0030
Fixed-source FM
96%
84.9
1.6
1.1
0.0038
CNP-Flow (Ours)
97%
88.5
0.8
0.5
0.0023
Table 3: Results on motion planning. All methods use a matched 15-step sampling budget, and results are averaged over three seeds.
Figure 4: Inference efficiency on KTH prediction. Curves show PSNR/SSIM averaged across generated samples per video. CNP-Flow reaches the standard Gaussian PSNR plateau ( ≈26.5 ) at NFE =5 and converges higher ( ≈28.2 ) by NFE =10 ; SSIM shows a similar trend. The crossover point is highlighted on each panel.
Figure 5: Qualitative comparison. Left : CLEVRER interpolation. Fixed-source FM drifts at the highlighted frames, while CNP-Flow matches the ground truth. Right : Franka Panda motion planning. Each panel overlays 50 sampled trajectories (red: colliding, green: collision-free). CNP-Flow produces more collision-free plans than MPD and fixed-source FM.
Video prediction (KTH)
Source x0
PSNR ↑
SSIM ↑
FVD ↓
N(0,I)
30.40
0.860
180.00
μϕ(Ci)
31.46
0.908
83.82
N(μϕ(Ci),σϕ2(Ci)I)
34.25
0.943
44.76
Table 4: Ablation studies. Left : source parameterization on KTH 10→30 video prediction, isolating the contribution of the learned mean and learned variance. Right : source learning and backbone fine-tuning (FT) on Franka Panda motion planning. In both tasks, metrics improve as the source better matches the condition.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Fixed-source FM
CNP-Flow
Latency increase
Additional memory
KTH / BAIR64
207.1 ms/frame
214.2 ms/frame
3.43%
∼15 MB
CLEVRER
211.9 ms/frame
219.1 ms/frame
3.39%
Appendix
Table 5: Inference overhead of CNP-Flow. Latency is measured over 1,000 generated frames under matched hardware, batch-size, and solver settings.
Component
KTH
CLEVRER
BAIR64
Flow reversal steps
100
20
50
FM fine-tune length
200k steps
20k steps
50k steps
FM fine-tune batch
64
128
64
FM fine-tune LR
5×10−5
5×10−5
5×10−5
Appendix
Table 6: Video prediction fixed-source FM and flow reversal configuration. Source samples obtained through flow reversal are cached only for CNP supervision and are not used at inference time.
Component
KTH
CLEVRER
FM pretrain steps
300k
300k
FM pretrain batch
32
16
Flow reversal steps
100
40
FM fine-tune steps
200k
200k
FM fine-tune batch
32
16
FM fine-tune LR
1×10−4
1×10−4
Appendix
Table 7: Video interpolation fixed-source FM and flow reversal configuration. Since no compatible interpolation checkpoint is available, the fixed-source CFM backbone is trained from scratch.
Component
Setting
Backbone
Convolutional encoder
Input
Conditional latents + temporal-distance embedding
Output
μ,logσ2
Base channels
64
Channel multipliers
(1,2,4,8)
Residual blocks per resolution
2
Appendix
Table 8: Video CNP architecture. The same architecture is used for video prediction and video interpolation.
Task
Dataset
Length
Batch
KL warm-up
Prediction
KTH
50 epochs
64
10 epochs
Prediction
CLEVRER
15 epochs
1024
10 epochs
Prediction
BAIR64
15 epochs
64
5 epochs
Interpolation
KTH
200k steps
32
5 epochs
Interpolation
CLEVRER
200k steps
16
2 epochs
Appendix
Table 9: Video CNP training configuration. All video CNPs are trained with Adam using learning rate 1×10−4 and KL weight λKL=1×10−4 .
Dataset
PSNR ↑
SSIM ↑
FVD ↓
Video prediction
BAIR64 1→15
20.15±0.13
0.8390±0.0021
61.17±1.58
KTH 10→30
34.25±0.26
0.9432±0.0025
44.76±1.53
KTH 10→40
32.69±0.22
0.9303±0.0040
54.88±3.85
CLEVRER 2→14
42.23±0.19
0.9935±0.0002
24.18±2.46
Video interpolation
Appendix
Table 10: Variability across runs for CNP-Flow. Results are reported as mean ± standard deviation over three independent runs for video prediction and interpolation.
Component
Setting
Backbone
Residual MLP
Input
Start state + EE goal pose + context embedding
Raw condition dimension
26
Context embedding dimension
128
Output
μ,logσ2
Hidden dimension
1024
Appendix
Table 11: Motion planning CNP architecture. The CNP predicts an isotropic Gaussian source distribution over the learnable B-spline control-point noise.
Method
Valid (%) ↑
EE pos. (cm) ↓
EE ori. ( ∘ ) ↓
Collision ↓
Diffusion prior
64.3±1.5
5.3±0.1
7.5±0.5
0.0402±0.0013
Fixed-source FM
61.1±3.8
3.8±0.2
5.6±0.5
0.0445±0.0034
CNP-Flow
64.3±4.6
2.1±0.2
2.9±0.3
0.0366±0.0045
Appendix
Table 12: Motion-planning performance without post-hoc refinement. Results are reported as mean ± standard deviation over three independent runs.
Regime
Task
Split
Method
MSE ↓
M-MSE ↓
PSNR ↑
SSIM ↑
Rigid
push_cube
IND
Vanilla FM
0.002886
0.017354
26.29
0.9546
CNP-Flow (Ours)
0.002192
0.013373
27.80
0.9647
OOD
Vanilla FM
0.002977
0.016039
26.78
0.9539
CNP-Flow (Ours)
0.002338
0.012326
28.13
0.9629
Deformable
push_rope
IND
Vanilla FM
0.000213
0.002708
36.75
0.9878
CNP-Flow (Ours)
0.000207
0.002608
36.88
0.9889
Appendix
Table 13: Results on the ACWM-Phys physical world-model benchmark at 240×240 resolution. We evaluate one task from each physical regime on both in-distribution (IND) and out-of-distribution (OOD) splits. All results are averaged over three independent seeds.
Source, frozen FM
PSNR ↑
SSIM ↑
FVD ↓
Standard Gaussian
30.29
0.8591
179.97
CNP w/o flow reversal supervision
30.46
0.8612
179.54
CNP w/ flow reversal supervision
30.85
0.9008
100.91
Appendix
Table 14: Effect of flow reversal supervision on source learning. The Stage-1 FM is frozen for all variants.
Source
Pairwise LPIPS ↑
Standard Gaussian
0.0666
CNP-Flow
0.0686
Appendix
Table 15: Prediction diversity on BAIR64. We report the average pairwise LPIPS among predictions generated from the same condition.
Dataset
Round 1
Round 2
Round 3
BAIR64
61.17
60.48
67.07
KTH
44.76
35.04
36.37
Appendix
Table 16: Iterative Stage-2/Stage-3 refinement. FVD is reported after each round of source learning and FM refinement.
Flow reversal solver
FVD ↓
Dopri
61.73±1.47
Heun-100
61.16±0.99
Euler-100
61.17±1.56
Appendix
Table 17: Robustness to the numerical solver used for flow reversal on BAIR64 prediction. Values are mean ± standard deviation over three seeds.
Method
∥x0−x1∥2↓
Path length ↓
Turning (rad) ↓
Endpoint dev. ↓
Standard Gaussian
19.1715
18.6780
0.8292
2.7362
CNP-Flow (Ours)
17.4091
17.0088
0.7586
2.4710
Appendix
Table 18: Transport geometry on BAIR64. Both methods are evaluated on the same 256 condition–target pairs using Euler-100.
Coupling
Distance ↓
Path ↓
Turning ↓
Endpoint dev. ↓
Flow reversal pair
15.7342
15.7185
0.2261
0.4913
Conditional OT
17.3715
16.2187
1.0610
3.9173
CNP-Flow (Ours)
17.4091
17.0088
0.7586
2.4710
Appendix
Table 19: Alternative Stage-3 couplings on BAIR64 prediction. Left : transport geometry on 256 condition–target pairs using Euler-100. Right : downstream generation quality. Direct pairing via flow reversal requires the target and is therefore reported only as a training-time geometry reference.
Figure 6: Toy example visualization of the flow reversal mapping and transport distance analysis.
In recent years Flow Matching has become a prominent method for generative modeling robot motion generation. In its generic form Flow Matching is an ODE-based neural sampler that is trained by regressing empirical flow fields associated with motion samples as data. However, in robot motion generation we often have additional constraints that might not be present in the collected data. The majority of current approaches train the flow on the available data and use inference-time guidance to enforce task-specific constraints. To address this mismatch, we propose \textbf{ConFlow}, a constraint-guided flow matching framework that incorporates constraint information directly into the training objective via differentiable barrier or cost functions. To address design specifications such as smoothness and boundary conditions, we propose replacing the standard Gaussian source distribution used in flow matching training with a conditional Gaussian Process. Our approach also uses infeasible demonstrations as negative supervision, improving constraint satisfaction without requiring additional expert data. Experiments on a two-robot navigation task demonstrate that ConFlow achieves lower collision rates and higher trajectory quality than standard flow matching baselines, with or without inference-time guidance. These results validate training-time constraint integration as an effective approach to closing the training--inference gap in generative motion models.
Nutan Chen, Jianxiang Feng, Marvin Alles +1
LS Wiiri Robot Innovation Center · Independent Researcher · Technical University of Munich +1
Flow Matching (FM) is a simulation-free method for learning a continuous, invertible flow that interpolates between two distributions, and in particular generates data from noise. Inspired by the variational nature of the diffusion process as a gradient flow, we introduce a stepwise FM model, Local Flow Matching (LFM), which sequentially learns a sequence of FM submodels, each matching a diffusion process up to the time-step size in the data-to-noise direction. In each step, the two distributions to be interpolated by the sub-flow model are closer than those in the full-flow matching model, which interpolates data to noise distributions, enabling smaller models with more efficient training. This variational perspective also allows us to prove a theoretical generation guarantee for the proposed flow model in terms of the χ2-divergence between the generated and true data distributions, leveraging the contraction property of the diffusion process. In practice, the stepwise structure of LFM is naturally amenable to model distillation, and various distillation techniques can be applied to accelerate generation. We empirically demonstrate that LFM achieves competitive generative performance compared to FM on unconditional generation of tabular and image datasets, and on conditional generation of robotic manipulation policies.
Chen Xu, Xiuyuan Cheng, Yao Xie
H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology. · Department of Mathematics, Duke University
We propose Coreset-Induced Conditional Velocity Flow Matching (CCVFM), a generative model that augments hierarchical rectified flow with a data-informed source distribution. Hierarchical flow matching models the full conditional velocity law in velocity space, but its inner flow is asked to transport isotropic Gaussian noise to a multimodal target velocity distribution from scratch. Our key observation is that this inner source can be replaced by a closed-form surrogate built from a coreset of the target. CCVFM first compresses the target into weighted atoms using an entropic Sinkhorn coreset and lifts them to a Gaussian mixture. The induced conditional velocity law is then a closed-form Gaussian mixture that can be sampled without a learned neural sampler. A lightweight correction flow, trained from this exact surrogate source, then refines the remaining surrogate-to-target residual rather than learning an entire noise-to-data map. We prove that the surrogate transport cost equals the target--surrogate Wasserstein gap under an explicit compression assumption, whereas the noise-source analogue has a dimension-scale lower bound. We further characterize the conditional second moment of the direct surrogate-source training target and show that its source-dependent excess is small when the surrogate conditional law is close to the true conditional velocity law in mean and covariance. Empirically, on MNIST, CIFAR-10, ImageNet-32, and CelebA-HQ, the proposed method reaches competitive few-step generation under matched architectures.