We study path-flow alignment as a unified training objective for flow matching. Instead of fixing the interpolation path and learning only the velocity field, we jointly train an endpoint-preserving path network and a flow network using the same alignment loss: the flow learns to match the path velocity, and the path learns to align its velocity to the current flow. Although every fixed learned path defines a valid flow-matching objective, the alignment loss alone is not a reliable criterion for path learning. We identify path overfitting, a failure mode in which the alignment loss decreases while sample quality worsens. We find that this failure is associated with low-entropy bottlenecks in the induced probability path, where the learned path routes samples through overly concentrated intermediate marginals. Motivated by this diagnosis, we introduce a stochastic path regularizer that hides part of the source information from the path network while preserving exact endpoints. The resulting regularization gives an explicit entropy floor for the stochastic training-path marginals and empirically suppresses the bottleneck in the learned sampler, making joint path-flow training effective. On ImageNet-256x256 with SiT backbones, our method consistently improves FID across model scales, extends to model-guidance training, and leaves the inference-time architecture and sampler unchanged. Code is available at https://github.com/lizeyu090312/traj_opt_paper
Figures & tables
Figure 1: Learned ODE trajectories in a toy mixture example. Top: standard flow matching, path–flow alignment without stochastic regularization ( ρ=1 ), and with stochastic regularization ( ρt=1−t ), shown at three times. Bottom: standard and unregularized trajectories overlaid. All methods start from the same source samples. Appendix F.4 gives further details.
Figure 2Figure 3
Objective
Model
Params
Init FID
Baseline best
Ours best
FID gain
FM
SiT-S
33M
9.109
9.022
8.359
7.3%
FM
SiT-B
130M
5.074
5.057
4.415
12.7%
FM
SiT-L
458M
3.424
3.390
2.742
19.1%
FM
SiT-XL
675M
2.091
2.091
1.825
12.7%
MG
SiT-B-MG
130M
8.090
7.940
5.585
29.7%
MG
SiT-XL-MG
675M
1.522
1.494
1.322
11.5%
Table 1: ImageNet- 256×256 summary. We report the best 50k-FID ( ↓ ) over the same 12k continuation budget; full checkpoint-wise FID/IS results are in Appendix Tables 3 – 4 . Params denote the inference-time SiT flow model only; the path network is used only during training.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
γlin
Standard linear path, γlin(t;x0,x1)=(1−t)x0+tx1 .
γ
Generic sufficiently regular path, possibly with auxiliary randomness ξ .
γψ
Unregularized learned residual path.
γ^ψ
Stochastic-regularized learned residual path.
pγ,t
Marginal distribution of a generic path γ(t;x0,x1,y,ξ) at time t .
pψ,t
Marginal distribution of the unregularized learned path γψ(t;x0,x1,y) .
Appendix
Table 2: Notation convention for paths, induced marginals, and velocity fields.
Figure 6: FID versus per-dimensional flow divergence in the last time bucket. The last bucket corresponds to the final bin of the ten-bin estimator in ( 11 ). Runs with large positive last-bucket divergence are precisely the high-FID overfit checkpoints in this sweep, while the best checkpoints keep the final-bin divergence negative or only mildly positive.
Figure 7: FID versus global flow-geometry summaries. Unlike the last-bucket divergence in Figure 6 , average flow speed and acceleration are not monotone indicators of sample quality. Some low-FID stochastic source-regularized checkpoints have comparable or larger average flow speed and acceleration than higher-FID unregularized checkpoints.
Figure 8: FID versus global path-geometry summaries. The path-speed and path-acceleration summaries also do not explain the FID ranking on their own. In particular, good stochastic source-regularized checkpoints can have larger path speed or acceleration than worse unregularized checkpoints. This supports the view that the key failure mode is distributional—a late entropy-expansion spike—rather than simply excessive path length or curvature.
Model
Params
FID (baseline) ↓
FID (ours) ↓
IS (baseline) ↑
IS (ours) ↑
Init
Best
4k
8k
12k
Init
4k
8k
12k
SiT-S
33M
9.109
9.022
8.375
8.359
8.492
173.6
188.9
187.2
185.4
SiT-B
130M
5.074
5.057
4.703
4.480
4.415
199.0
234.9
229.0
225.5
SiT-L
458M
3.424
3.390
3.014
2.882
2.742
201.1
274.8
271.5
267.2
SiT-XL
675M
2.091
2.091
1.909
1.845
1.825
256.1
282.9
285.6
288.6
Appendix
Table 3: Full ImageNet- 256×256 results corresponding to Table 1 . The baseline columns report the initial checkpoint and the best FID reached by standard fine-tuning within the 12k continuation budget. Our method is evaluated after each alternating cycle.
Model
Method
0k
4k
8k
12k
SiT-B-MG (130M)
Ours
8.090
6.704
5.844
5.585
MG baseline
8.090
8.011
7.940
7.950
SiT-XL-MG (675M)
Ours
1.522
1.336
1.322
1.333
MG baseline
1.522
1.494
1.504
1.501
Appendix
Table 4: Checkpoint-wise model-guidance FID results corresponding to Table 1 . Both methods start from the same model-guidance checkpoint; lower is better.
Model
Alg 1 (Path Phase)
Alg 1 (Flow Phase)
Standard flow matching
SiT-S
1.86
1.39
4.65
SiT-B
1.70
1.02
2.08
SiT-L
0.85
0.51
0.71
SiT-XL
0.70
0.37
0.50
Appendix
Table 5: Training efficiency measured in steps per second per H200 GPU.
Model
Ours GPU-h (path + flow)
Matched baseline steps
Baseline GPU-h
SiT-S
3.574 (1.210 + 2.364)
61.1k
3.579
SiT-B
4.667 (1.468 + 3.199)
36.6k
4.677
SiT-XL
12.216 (3.546 + 8.670)
22.8k
12.243
Appendix
Table 6: GPU-hour matching for three training cycles. GPU-hours are measured on H200 GPUs.
Model
Initialization
Baseline best
Compute-matched
Ours
Reduction
SiT-S
9.109
9.022
9.025
8.359
7.3%
SiT-B
5.074
5.057
4.981
4.415
11.4%
SiT-XL
2.091
2.091
2.147
1.825
12.7%
Appendix
Table 7: Compute-matched ImageNet- 256×256 continuation. FID is computed from 50k images using the Heun sampler at 250 NFE. The reduction is relative to the stronger standard baseline.
SiT-S
CFG 2
CFG 3
CFG 3.5
CFG 4
Baseline
18.742
13.094
13.099
13.967
Ours
20.364
11.463
10.988
11.439
Appendix
Table 8: FID on the 10k-image CFG-selection sweeps. Bold and underline identify the selected value for each checkpoint.
SiT-S
Precision ↑
Recall ↑
Baseline
0.874
0.283
Ours
0.860
0.297
Appendix
Table 9: Precision and recall corresponding to the 50k-image headline FID results in Table 1 , generated using Heun at 250 NFE. Bold marks the better value within each model and metric.
SiT-S
Metric
Method
CFG 2
CFG 3
CFG 3.5
CFG 4
Precision ↑
Baseline
0.850
0.875
0.872
0.868
Ours
0.831
0.858
0.857
0.852
Recall ↑
Baseline
0.360
0.364
0.354
0.332
Ours
0.385
0.433
0.416
0.407
SiT-B
Appendix
Table 10: Precision and recall across 10k-image CFG sweeps using the Dopri5 sampler.
Model
Baseline best
Ours best
SiT-S
9.142±0.13
8.383±0.041
SiT-B
5.093±0.041
4.470±0.054
SiT-L
3.458±0.061
2.777±0.031
SiT-XL
2.103±0.026
1.830±0.015
Appendix
Table 11: Mean ± sample standard deviation of 50k-image FID across three generation-seed ranges, using Heun at 250 NFE.
Checkpoint
CFG 1.5
CFG 2.0
CFG 2.5
Baseline
11.890
9.142
11.230
Ours
15.734
8.204
8.765
Appendix
Table 12: FlowDCN-B 10k-image CFG sweep using the 100-step adaptive Dopri5 sampler.
Model
Params
Init FID
Baseline best
Ours best
FID gain
FlowDCN-B
120M
6.489
5.888
5.473
7.0%
Appendix
Table 13: Performance of path–flow alignment using FlowDCN-B (Ours best) compared to continued finetuning (Baseline best), 50k-FID (Heun sampler 250 NFE).
Method
FID, 2 NFE ↓
FID, 40 NFE ↓
Our method (1 cycle)
170.845
9.097
Baseline continuation
100.518 (step 1,000)
10.339 (step 1,500)
[ 17 ] , λmax=0.001 , GLOW LR 5×10−6
79.150
11.667
[ 17 ] , λmax=0.001 , GLOW LR 5×10−7
84.230
11.568
[ 17 ] , λmax=0.0003 , GLOW LR 5×10−7
94.295
11.878
Appendix
Table 14: FLOP-matched CIFAR-10 comparison using 10k images from one generation seed. Lower FID is better.
Figure 9: Uncurated SiT-XL samples matched by class and generation seed. Left: the best baseline, given by the framework initialization. Right: path–flow alignment after three cycles.
Recent work has popularized a practical recipe in diffusion and flow matching: predict the clean signal x, convert it to a velocity, and train through a velocity-space loss. The conversion contains a singular endpoint amplification and therefore appears prone to unstable optimization, yet recent systems obtain strong empirical results with this recipe. We investigate this tension through the integrability of the pre-optimizer stochastic-gradient second moment. Under stated initialization conditions, the moment diverges under Uniform sampling; boundary-suppressing sampling can restore integrability under an additional upper-growth condition. We then show that prediction--loss alignment eliminates this conversion-induced source of non-integrability. Under a uniform moment bound, alignment yields a finite second moment for every timestep density, including Uniform sampling. Controlled experiments across continuous and binary settings reproduce the predicted sampler-dependent instability and show that aligned objectives remain trainable across the tested samplers. These results reconcile pointwise amplification with sampler-dependent empirical success and support alignment as a principled route to more robust flow-matching training.
Jiadong Hong, Lei Liu, Xinyu Bian +2
Zhejiang University, Hangzhou, China · Huawei Technologies Co., Ltd, Hong Kong
Trajectory crossing remains a critical bottleneck in Flow Matching (FM), and previous works typically view these crossings from a theoretical optimization perspective causing velocity averaging. They attempt to address it indirectly by post-hoc distillation or endpoint coupling, without explicitly regulating the intermediate trajectories. In this paper, we introduce a new network learning perspective: crossing points inherently induce large local Lipschitz constants in the target velocity field, leading to two drawbacks. First, high Lipschitz constants correspond to high-frequency signals in the velocity field that neural networks struggle to fit due to spectral bias. Second, they also imply drastic velocity variations, leading to severe numerical integration errors in few-step inference. To alleviate this, we propose CoFlow, a framework that introduces the contrastive learning paradigm into FM to explicitly repel trajectories during training, thereby lowering the local Lipschitz constants of the velocity field. Specifically, we formulate CoFlow from a Stochastic Differential Equation (SDE) perspective by injecting a repulsive drift term. This drift actively guides the forward process of positive samples away from negative trajectories, effectively reducing the local Lipschitz constant. Furthermore, we derive an equivalent stochastic interpolant formulation from this SDE, providing a simple and tractable design space to control the influence of negative samples. Extensive experiments on ImageNet 256x256 demonstrate that CoFlow significantly reduces FID compared to standard FM in few-step inference (e.g., 20 steps), with no added training overhead. The code can be accessed at: https://github.com/HKUST-LongGroup/CoFlow
Ziqi Jiang, Zhenqi He, Long Chen
Department of CSE, The Hong Kong University of Science and Technology
Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-dimensional images, each training sample supervises only a single trajectory and intermediate point, yielding an extremely sparse and high-variance training signal. This under-constrained supervision can cause flow collapse, where the learned dynamics memorize specific source-target pairings, mapping diverse inputs to overly similar outputs, failing to generalize. We introduce Posterior-Augmented Flow Matching (PAFM), a theoretically grounded generalization of FM that replaces single-target supervision with an expectation over an approximate posterior of valid target completions for a given intermediate state and condition. PAFM factorizes this intractable posterior into (i) the likelihood of the intermediate under a hypothesized endpoint and (ii) the prior probability of that endpoint under the condition, and uses an importance sampling scheme to construct a mixture over multiple candidate targets. We prove that PAFM yields an unbiased estimator of the original FM objective while substantially reducing gradient variance during training by aggregating information from many plausible continuation trajectories per intermediate. Finally, we show that PAFM improves over FM by up to 3.4 FID50K across different model scales (SiT-B/2 and SiT-XL/2), different architectures (SiT and MMDiT), and in both class and text conditioned benchmarks (ImageNet and CC12M), with a negligible increase in the compute overhead. Code: https://github.com/gstoica27/PAFM.git.
George Stoica, Sayak Paul, Matthew Wallingford +6
Georgia Tech · University of Washington · Hugging Face +2