Flow matching (FM) learns generative transport by fitting continuous-time motion from a simple source distribution to the data distribution. Most existing methods use first-order bridges: once a source and a target sample are paired, the path is a straight motion with constant velocity. FM with optimal transport (OT) improves the pairing, but the bridge itself remains linear, limiting its ability to model curved motion, acceleration, and changing directions. A natural remedy is to use second-order phase-space dynamics; however, learning the bridge requires target-side terminal-velocity information that static datasets do not provide. We propose Variational Terminal-Velocity Flow Matching (VTV-FM), a second-order FM framework that derives the missing velocity by minimizing acceleration energy, yielding a closed-form closure for static data. The same minimum-acceleration variational construction also defines the OT pairing cost and the acceleration targets used for training. Experiments on low-dimensional datasets, PDE-governed physical fields, and CIFAR-10 show that VTV-FM improves transport geometry and generation quality over first-order and high-order FM baselines.
Figures & tables
Figure 1: Effect of terminal-velocity choices on second-order bridge supervision. All curves in (a,b) use the same observable boundary information (x0,v0,x1) , but different terminal velocities v1 . In (a,b), the horizontal axis is normalized time t/T , and colors denote the same terminal-velocity choices: v1⋆ , v1=0 , chord velocity (x1−x0)/T , random velocity, and v1=v0 . (a) The position bridges show the evolution of x(t) under different terminal-velocity choices. (b) The velocity bridges show the corresponding evolution of v(t) . (c) The bars show the acceleration-energy cost for each terminal-velocity choice. Lower cost corresponds to smoother supervision, while externally specified terminal velocities may induce biased supervision.
Table 1: Low-dimensional results measured by sample-based W2 . Lower is better. The right radar plot visualizes the same results using the relative score sd,m , where higher is better.
Dataset
Method
MMD
PSD log- L2
Deriv. gap
NS2D
OT-CFM
0.017
5.08
6.93×10−3
HOFM
0.015
4.22
1.06×10−3
VTV-FM
0.014
3.48
2.80×10−4
KS2D
OT-CFM
0.035
9.00
5.50×10−1
HOFM
0.015
8.13
9.17×10−2
VTV-FM
0.010
7.52
1.30×10−2
Table 2: PDE generation results.
Method
Euler-100
Dopri5
Heun-100
Taylor2-100
Verlet-100
FM
4.58
3.73 (147)
3.62
–
–
I-CFM
4.48
3.68 (146)
3.59
–
–
OT-CFM
4.46
3.60 (134)
3.48
–
–
FM + OAT-FM
3.91
3.56 (135)
3.51
–
–
I-CFM + OAT-FM
3.77
3.50 (138)
3.45
–
–
OT-CFM + OAT-FM
3.72
3.47 (126)
3.42
–
–
Table 3: CIFAR-10 FID under controlled inference protocols. For adaptive Dopri5, the measured number of function evaluations (NFE) is shown in parentheses.
Objective
Verlet
Heun
Taylor2
Acc.
3.10
3.17
3.25
Acc. + Cons.
3.06
3.16
3.23
Acc. + ADM
3.05
3.14
3.19
Acc. + Cons. + ADM
2.99
3.13
3.18
Table 4: VTV-FM objective ablation.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Transport trajectories on six low-dimensional datasets. Each panel compares CFM , OT-CFM , HOFM , and VTV-FM .
Figure 3: NS2D and KS2D generation examples. Within each example, the top, middle, and bottom rows show real, VTV-FM, and OT-CFM samples, respectively.
Figure 4: Qualitative CIFAR-10 samples generated by VTV-FM.
Figure 5: Qualitative temporal examples from KTH, BAIR, and Navier–Stokes sequences.
Loss ablation
Solver comparison
Dataset
Acc.
+Cons.
+ADM
Vel.+Acc.
Heun
Taylor2
Verlet
Dinosaur
0.1119
0.1114
0.1112
0.1109
0.1113
0.1121
0.1109
Swiss
0.1112
0.1081
0.1105
0.1127
0.1081
0.1092
0.1105
Checker
0.1092
0.1094
0.1098
0.1092
0.1092
0.1096
0.1099
Moons
0.0835
0.0826
0.0807
0.0790
0.0790
0.0803
0.0808
Lines
0.1362
0.1366
0.1380
0.1387
0.1370
0.1362
0.1376
Appendix
Table 5: Ablation and fixed-step solver analysis for VTV-FM on low-dimensional datasets. Left: loss variants. Right: solver comparison under the best-performing loss setting for each dataset.
σ
0
0.5
1.0
1.5
2.0
Avg. W2
0.1283
0.1002
0.1004
0.1011
0.1056
Appendix
Table 6: Sensitivity to the source-velocity scale, averaged over the seven low-dimensional datasets.
Setting
v1⋆
Zero
Chord
Random
Copy- v0
Tangent
Fixed pairing
0.1004
0.1050
0.1020
0.1094
0.1062
0.1013
End-to-end
0.1004
0.1041
0.1008
0.1087
0.1054
0.1005
Appendix
Table 7: Terminal-velocity ablation averaged over the seven low-dimensional datasets.
Method
Bridge
Pairing
Avg. W2
OT-CFM
Linear
Standard OT
0.1077
2OT-CFM
Linear
Projected 2OT
0.1109
OT-VTV-FM
Second-order
Standard OT
0.1283
2OT-VTV-FM
Second-order
Projected 2OT
0.1004
Appendix
Table 8: Pairing–bridge cross-ablation on the low-dimensional datasets.
Solver
Method
Avg. W2
NFE
Time (ms)
Peak memory (MB)
Heun-100
CFM
0.1085
200
858
62.30
OT-CFM
0.1079
200
861
62.30
VTV-FM
0.1020
200
859
62.37
Dopri5
CFM
0.1069
72.8
486
62.58
OT-CFM
0.1066
42.2
306
62.59
VTV-FM
0.1018
38.0
242
63.66
Appendix
Table 9: Compute-matched comparison on the low-dimensional benchmarks.
Dataset
VTV-FM
VTV-FM-Jerk-ZA
Dinosaur
0.1119
0.1113
Swiss
0.1112
0.1120
Checker
0.1092
0.1096
Moons
0.0835
0.0829
Lines
0.1362
0.1356
Tree
0.0741
0.0745
Appendix
Table 10: Bridge-shape ablation using acceleration matching only. VTV-FM-Jerk-ZA uses a minimum-jerk bridge with zero endpoint accelerations.
Department of Mechanical Engineering Colorado State University Fort Collins, CO 80523, USA · Department of Electrical and Computer Engineering Colorado State University Fort Collins, CO 80523, USA