We study path-flow alignment as a unified training objective for flow matching. Instead of fixing the interpolation path and learning only the velocity field, we jointly train an endpoint-preserving path network and a flow network using the same alignment loss: the flow learns to match the path velocity, and the path learns to align its velocity to the current flow. Although every fixed learned path defines a valid flow-matching objective, the alignment loss alone is not a reliable criterion for path learning. We identify path overfitting, a failure mode in which the alignment loss decreases while sample quality worsens. We find that this failure is associated with low-entropy bottlenecks in the induced probability path, where the learned path routes samples through overly concentrated intermediate marginals. Motivated by this diagnosis, we introduce a stochastic path regularizer that hides part of the source information from the path network while preserving exact endpoints. The resulting regularization gives an explicit entropy floor for the stochastic training-path marginals and empirically suppresses the bottleneck in the learned sampler, making joint path-flow training effective. On ImageNet-256x256 with SiT backbones, our method consistently improves FID across model scales, extends to model-guidance training, and leaves the inference-time architecture and sampler unchanged. Code is available at https://github.com/lizeyu090312/traj_opt_paper
Figures & tables
Figure 1: Learned ODE trajectories in a toy mixture example. Top: standard flow matching, path–flow alignment without stochastic regularization ( ρ=1 ), and with stochastic regularization ( ρt=1−t ), shown at three times. Bottom: standard and unregularized trajectories overlaid. All methods start from the same source samples. Appendix F.4 gives further details.
Figure 2Figure 3
Objective
Model
Params
Init FID
Baseline best
Ours best
FID gain
FM
SiT-S
33M
9.109
9.022
8.359
7.3%
FM
SiT-B
130M
5.074
5.057
4.415
12.7%
FM
SiT-L
458M
3.424
3.390
2.742
19.1%
FM
SiT-XL
675M
2.091
2.091
1.825
12.7%
MG
SiT-B-MG
130M
8.090
7.940
5.585
29.7%
MG
SiT-XL-MG
675M
1.522
1.494
1.322
11.5%
Table 1: ImageNet- 256×256 summary. We report the best 50k-FID ( ↓ ) over the same 12k continuation budget; full checkpoint-wise FID/IS results are in Appendix Tables 3 – 4 . Params denote the inference-time SiT flow model only; the path network is used only during training.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
γlin
Standard linear path, γlin(t;x0,x1)=(1−t)x0+tx1 .
γ
Generic sufficiently regular path, possibly with auxiliary randomness ξ .
γψ
Unregularized learned residual path.
γ^ψ
Stochastic-regularized learned residual path.
pγ,t
Marginal distribution of a generic path γ(t;x0,x1,y,ξ) at time t .
pψ,t
Marginal distribution of the unregularized learned path γψ(t;x0,x1,y) .
Appendix
Table 2: Notation convention for paths, induced marginals, and velocity fields.
Figure 6: FID versus per-dimensional flow divergence in the last time bucket. The last bucket corresponds to the final bin of the ten-bin estimator in ( 11 ). Runs with large positive last-bucket divergence are precisely the high-FID overfit checkpoints in this sweep, while the best checkpoints keep the final-bin divergence negative or only mildly positive.
Figure 7: FID versus global flow-geometry summaries. Unlike the last-bucket divergence in Figure 6 , average flow speed and acceleration are not monotone indicators of sample quality. Some low-FID stochastic source-regularized checkpoints have comparable or larger average flow speed and acceleration than higher-FID unregularized checkpoints.
Figure 8: FID versus global path-geometry summaries. The path-speed and path-acceleration summaries also do not explain the FID ranking on their own. In particular, good stochastic source-regularized checkpoints can have larger path speed or acceleration than worse unregularized checkpoints. This supports the view that the key failure mode is distributional—a late entropy-expansion spike—rather than simply excessive path length or curvature.
Model
Params
FID (baseline) ↓
FID (ours) ↓
IS (baseline) ↑
IS (ours) ↑
Init
Best
4k
8k
12k
Init
4k
8k
12k
SiT-S
33M
9.109
9.022
8.375
8.359
8.492
173.6
188.9
187.2
185.4
SiT-B
130M
5.074
5.057
4.703
4.480
4.415
199.0
234.9
229.0
225.5
SiT-L
458M
3.424
3.390
3.014
2.882
2.742
201.1
274.8
271.5
267.2
SiT-XL
675M
2.091
2.091
1.909
1.845
1.825
256.1
282.9
285.6
288.6
Appendix
Table 3: Full ImageNet- 256×256 results corresponding to Table 1 . The baseline columns report the initial checkpoint and the best FID reached by standard fine-tuning within the 12k continuation budget. Our method is evaluated after each alternating cycle.
Model
Method
0k
4k
8k
12k
SiT-B-MG (130M)
Ours
8.090
6.704
5.844
5.585
MG baseline
8.090
8.011
7.940
7.950
SiT-XL-MG (675M)
Ours
1.522
1.336
1.322
1.333
MG baseline
1.522
1.494
1.504
1.501
Appendix
Table 4: Checkpoint-wise model-guidance FID results corresponding to Table 1 . Both methods start from the same model-guidance checkpoint; lower is better.
Model
Alg 1 (Path Phase)
Alg 1 (Flow Phase)
Standard flow matching
SiT-S
1.86
1.39
4.65
SiT-B
1.70
1.02
2.08
SiT-L
0.85
0.51
0.71
SiT-XL
0.70
0.37
0.50
Appendix
Table 5: Training efficiency measured in steps per second per H200 GPU.
Model
Ours GPU-h (path + flow)
Matched baseline steps
Baseline GPU-h
SiT-S
3.574 (1.210 + 2.364)
61.1k
3.579
SiT-B
4.667 (1.468 + 3.199)
36.6k
4.677
SiT-XL
12.216 (3.546 + 8.670)
22.8k
12.243
Appendix
Table 6: GPU-hour matching for three training cycles. GPU-hours are measured on H200 GPUs.
Model
Initialization
Baseline best
Compute-matched
Ours
Reduction
SiT-S
9.109
9.022
9.025
8.359
7.3%
SiT-B
5.074
5.057
4.981
4.415
11.4%
SiT-XL
2.091
2.091
2.147
1.825
12.7%
Appendix
Table 7: Compute-matched ImageNet- 256×256 continuation. FID is computed from 50k images using the Heun sampler at 250 NFE. The reduction is relative to the stronger standard baseline.
SiT-S
CFG 2
CFG 3
CFG 3.5
CFG 4
Baseline
18.742
13.094
13.099
13.967
Ours
20.364
11.463
10.988
11.439
Appendix
Table 8: FID on the 10k-image CFG-selection sweeps. Bold and underline identify the selected value for each checkpoint.
SiT-S
Precision ↑
Recall ↑
Baseline
0.874
0.283
Ours
0.860
0.297
Appendix
Table 9: Precision and recall corresponding to the 50k-image headline FID results in Table 1 , generated using Heun at 250 NFE. Bold marks the better value within each model and metric.
SiT-S
Metric
Method
CFG 2
CFG 3
CFG 3.5
CFG 4
Precision ↑
Baseline
0.850
0.875
0.872
0.868
Ours
0.831
0.858
0.857
0.852
Recall ↑
Baseline
0.360
0.364
0.354
0.332
Ours
0.385
0.433
0.416
0.407
SiT-B
Appendix
Table 10: Precision and recall across 10k-image CFG sweeps using the Dopri5 sampler.
Model
Baseline best
Ours best
SiT-S
9.142±0.13
8.383±0.041
SiT-B
5.093±0.041
4.470±0.054
SiT-L
3.458±0.061
2.777±0.031
SiT-XL
2.103±0.026
1.830±0.015
Appendix
Table 11: Mean ± sample standard deviation of 50k-image FID across three generation-seed ranges, using Heun at 250 NFE.
Checkpoint
CFG 1.5
CFG 2.0
CFG 2.5
Baseline
11.890
9.142
11.230
Ours
15.734
8.204
8.765
Appendix
Table 12: FlowDCN-B 10k-image CFG sweep using the 100-step adaptive Dopri5 sampler.
Model
Params
Init FID
Baseline best
Ours best
FID gain
FlowDCN-B
120M
6.489
5.888
5.473
7.0%
Appendix
Table 13: Performance of path–flow alignment using FlowDCN-B (Ours best) compared to continued finetuning (Baseline best), 50k-FID (Heun sampler 250 NFE).
Method
FID, 2 NFE ↓
FID, 40 NFE ↓
Our method (1 cycle)
170.845
9.097
Baseline continuation
100.518 (step 1,000)
10.339 (step 1,500)
[ 17 ] , λmax=0.001 , GLOW LR 5×10−6
79.150
11.667
[ 17 ] , λmax=0.001 , GLOW LR 5×10−7
84.230
11.568
[ 17 ] , λmax=0.0003 , GLOW LR 5×10−7
94.295
11.878
Appendix
Table 14: FLOP-matched CIFAR-10 comparison using 10k images from one generation seed. Lower FID is better.
Figure 9: Uncurated SiT-XL samples matched by class and generation seed. Left: the best baseline, given by the framework initialization. Right: path–flow alignment after three cycles.