Trajectory crossing remains a critical bottleneck in Flow Matching (FM), and previous works typically view these crossings from a theoretical optimization perspective causing velocity averaging. They attempt to address it indirectly by post-hoc distillation or endpoint coupling, without explicitly regulating the intermediate trajectories. In this paper, we introduce a new network learning perspective: crossing points inherently induce large local Lipschitz constants in the target velocity field, leading to two drawbacks. First, high Lipschitz constants correspond to high-frequency signals in the velocity field that neural networks struggle to fit due to spectral bias. Second, they also imply drastic velocity variations, leading to severe numerical integration errors in few-step inference. To alleviate this, we propose CoFlow, a framework that introduces the contrastive learning paradigm into FM to explicitly repel trajectories during training, thereby lowering the local Lipschitz constants of the velocity field. Specifically, we formulate CoFlow from a Stochastic Differential Equation (SDE) perspective by injecting a repulsive drift term. This drift actively guides the forward process of positive samples away from negative trajectories, effectively reducing the local Lipschitz constant. Furthermore, we derive an equivalent stochastic interpolant formulation from this SDE, providing a simple and tractable design space to control the influence of negative samples. Extensive experiments on ImageNet 256x256 demonstrate that CoFlow significantly reduces FID compared to standard FM in few-step inference (e.g., 20 steps), with no added training overhead. The code can be accessed at: https://github.com/HKUST-LongGroup/CoFlow
Figures & tables
Figure 1: Illustration of our motivation. (a) In standard FM, straight trajectories inevitably cross. At the intersection, spatially adjacent points have conflicting velocity targets ( vA=vB while xA≈xB ), leading to an extremely high local Lipschitz constant L . (b) CoFlow explicitly repels trajectories to alleviate crossing, keeping xA and xB separated. This significantly reduces the local Lipschitz constant, yielding a smoother target velocity field that is easier to learn and induces lower numerical integration errors during few-step sampling.
Figure 2: Comparison between DeltaFM and CoFlow.
Figure 3: 2D Example: Gaussian to Two Moons. (a) Visualization of generated trajectories: CoFlow yields paths with reduced curvature compared to standard FM. (b) Local Lipschitz constant of the learned velocity field. (c) Relative Truncation Error reduced compared with standard FM.
Figure 4: Qualitative comparison . All images are generated using a 20-step EM solver.
Model
Method
FID ↓
sFID ↓
IS ↑
Precision ↑
Recall ↑
SiT-B/2
FM
61.43
35.18
27.46
0.43
0.54
CoFlow
52.79
21.71
35.00
0.45
0.58
FM + REPA
47.64
34.33
40.75
0.48
0.59
CoFlow + REPA
37.76
20.23
52.92
0.51
0.60
SiT-XL/2
FM
38.50
29.27
50.14
0.54
0.58
CoFlow
29.65
17.41
64.12
0.57
0.59
Table 1: Class-Conditional ImageNet 256 × 256. Training for 400k steps. We sample 50K images using the EMSolver, with the number of function evaluations (NFE) set to 20. No CFG involved.
Figure 6
λ
Negative signal x−
Sampling policy
FID ↓
sFID ↓
IS ↑
Standard FM
86.99
43.47
16.25
0.2
ϵ
Random
86.49
41.77
16.49
0.2
xtr−xt
Random
79.55
26.63
20.36
0.2
x1r−x1
Random
92.24
47.22
16.25
0.2
vtr−vt
Random
122.19
87.04
9.41
0.2
xtr−xt
Random
79.55
26.63
20.36
Table 4: Component-wise ablation on ImageNet 256 × 256. All models were SiT-S/2 trained for 400K iterations without REPA. We evaluated 50K images using the EM sampler with NFE =20 and fixed β(t)=t(1−t) . Colored cells indicate the factor varied in each block.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Model
Method
FID ↓
sFID ↓
IS ↑
SiT-B/2
FM
27.64
11.76
61.07
DeltaFM
20.40
5.48
70.41
FM + VC
19.97
5.45
71.44
SiT-XL/2
FM
11.40
8.69
114.69
DeltaFM
7.29
4.93
129.89
FM + VC
7.38
4.84
130.25
Appendix
Table 5: C omparison with DeltaFM. Both models are trained with REPA involved and NEF=50 for inference. “VC” denotes the inference-time Velocity Centering operation.
Flow matching trains a neural network to regress the conditional velocity along a linear interpolant between noise and data, and the number of network evaluations~(NFE) sets the cost of sampling. The straight-line interpolant carries an implicit choice: the sample moves at constant speed throughout the trajectory. We relax this choice and introduce Velocity Scheduled Flow Matching~(VSFM), which replaces the conditional target x1−x0 with v(t)(x1−x0) for any nonnegative profile v:[0,1]→R≥0 satisfying ∫01vdt=1. We study six polynomial profiles drawn from motion planning. The first use of VSFM is at inference time: a pretrained linear flow-matching model can be sampled under any admissible profile by integrating its ODE on a non-uniform τ-schedule, with no retraining and no additional computation; on CIFAR-10 this lowers FID by up to 19.8%. Training from scratch under a braking profile gives a further reduction of 17.4% at 4~NFE. Both gains follow from the local truncation error of the Euler integrator on the induced grid.
Vitalii Bondar
Cherkasy State Technological University Cherkasy, Ukraine
Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-dimensional images, each training sample supervises only a single trajectory and intermediate point, yielding an extremely sparse and high-variance training signal. This under-constrained supervision can cause flow collapse, where the learned dynamics memorize specific source-target pairings, mapping diverse inputs to overly similar outputs, failing to generalize. We introduce Posterior-Augmented Flow Matching (PAFM), a theoretically grounded generalization of FM that replaces single-target supervision with an expectation over an approximate posterior of valid target completions for a given intermediate state and condition. PAFM factorizes this intractable posterior into (i) the likelihood of the intermediate under a hypothesized endpoint and (ii) the prior probability of that endpoint under the condition, and uses an importance sampling scheme to construct a mixture over multiple candidate targets. We prove that PAFM yields an unbiased estimator of the original FM objective while substantially reducing gradient variance during training by aggregating information from many plausible continuation trajectories per intermediate. Finally, we show that PAFM improves over FM by up to 3.4 FID50K across different model scales (SiT-B/2 and SiT-XL/2), different architectures (SiT and MMDiT), and in both class and text conditioned benchmarks (ImageNet and CC12M), with a negligible increase in the compute overhead. Code: https://github.com/gstoica27/PAFM.git.
George Stoica, Sayak Paul, Matthew Wallingford +6
Georgia Tech · University of Washington · Hugging Face +2
While Flow Matching theoretically guarantees constant-velocity trajectories, we identify a critical breakdown in high-dimensional practice: the Velocity Deficit. We show that the MSE objective systematically underestimates velocity magnitude, causing generated samples to fail to reach the data manifold-a phenomenon we term Integration Lag. To rectify this, we propose Initial Energy Injection, instantiated via two complementary methods: the training-based Magnitude-Aware Flow Matching (MAFM) and the training-free Scale Schedule Corrector (SSC). Both are grounded in our discovery of a crucial asymmetry: velocity contraction causes harmful kinetic stagnation at the trajectory's start, yet acts as a beneficial denoising mechanism at its end. Empirically, SSC yields significant efficiency gains with zero retraining and just one line of code. On ImageNet-1k (256x256), it improves FID by 44.6% (from 13.68 to 7.58) and achieves a 5x speedup, enabling a 50-step generator (FID 7.58) to beat a 250-step baseline (FID 8.65). Furthermore, our methods generalize to Text-to-Image tasks and high-resolution generation, improving FID on MS-COCO by ~22%.