cs.LGSep 28, 2026

What Does a Stream Model Buy You in Flow Matching?

Authors: Jian Xu

Organizations: RIKEN

Abstract

Stream-level flow matching replaces the linear interpolant of conditional flow matching (CFM) by a Gaussian-process (GP) stream connecting each source--target pair, and reports lower sample error than \icfm{} on 2-Gaussian, MNIST and CIFAR-10 benchmarks. We ask what such a stream model actually contributes. Three results answer the question. (i)\emph{Reduction.} The stream-level CFM objective depends on the stream law only through the per-time joint law of (st,⋅t)(s_t,\sdot_t), so the conditional paths a Gaussian stream can reach are exactly the Gaussian conditional paths CFM already parametrises; in the coordinate-wise, shared-scalar-kernel construction gpcfm actually uses, the entire design space collapses to two scalar curves (mt,vt)(m_t,v_t), and cross-time covariance affects only estimator variance. (ii)\emph{The GP is a constrained chart of that space.} One kernel sets both mtm_t and vtv_t, so the paper's own recipe for widening coverage-shrinking the SE length-scale---destroys the interpolant (the midpoint mean weight falls from 1.031.03 to 0.000.00). On the 2-Gaussian benchmark this makes the GP chart diverge on 15/20015/200 runs at high coverage against 0/2000/200 for a decoupled (mt,vt)(m_t,v_t) chart (p=6.6×10−5p=6.6\times10^{-5}), and crossing the two curves shows the divergence tracks the mean, not the variance. On MNIST the same sweep does not diverge and the ordering reverses, so whether the coupling is harmful is benchmark-dependent; what holds on both is that the recipe buys nothing---no coverage level beats the paper's own, and past max⁡tvt≈0.6\max_t\sqrt{v_t}\approx0.6 both charts degrade. (iii)~\emph{Audit.} The released code does not implement the mechanism it describes: state and velocity are drawn independently (corr=0.00±0.01\mathrm{corr}=0.00\pm0.01 against an intended ±0.83\pm0.83--0.990.99).

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 15, 2026cs.LG

Same Flow, Different Paths: Variance Reduction in Flow Matching

In flow matching (FM), a velocity model vθv_θ is trained using a predefined path gtg_t that connects data and noise samples (e.g., gt(x0,x1)=(1−t)x0+tx1g_t(x_0, x_1) = (1 - t) x_0 + t x_1). In this work, we study the choice of this path from an optimization perspective by analyzing the variance of stochastic gradients. We consider the class G(pt,vt⋆)G(p_t,v^\star_t) of paths that induce the same marginal distributions ptp_t and marginal velocity field vt⋆v^\star_t, and therefore the same FM objective. Our main finding is that the choice of path gtg_t can fundamentally change the convergence rate of SGD, even when the FM objective remains exactly the same. (i) For a linear velocity model and one-dimensional Gaussian data, we derive a tight bound on the SGD iteration complexity up to logarithmic factors and find an analytically optimal path that minimizes this bound among linear paths inducing the same FM problem. (ii) We then extend the variance analysis to general FM problems and formulate path selection at a fixed θθ as the variance-minimization problem PathOptθ_θ, constrained to gt∈G(pt,vt⋆)g_t\in G(p_t,v^\star_t). We show that this constraint is essential: reducing variance without it can lead to slower convergence. (iii) Since the constraint gt∈G(pt,vt⋆)g_t \in G(p_t,v^\star_t) cannot generally be verified directly, we derive an equivalent formulation with constraints that can be estimated from samples, allowing paths to be found numerically. Our theoretical results are supported by experiments with Gaussian data, Gaussian mixture models, and real datasets.
Sep 30, 2026cs.LG

Correcting CondOT: Exact Finite-Step Sampling in Gaussian Flow Matching

Flow matching generates samples by gradually transforming noise into data. In practice, using a finite number of sampling steps introduces a numerical error that depends on the chosen schedule. We study this dependence for Gaussian targets and the explicit midpoint sampling method, using the exact flow field. We measure sampling error by the squared Wasserstein distance between the target distribution and the final distribution produced by the midpoint sampler. We show that the standard conditional optimal transport (CondOT) schedule cancels the leading midpoint error and improves the general convergence bound, even when the sampling steps are unequally spaced. On a uniform grid of SS sampling steps, we fix the signal schedule at αt=tα_t=t and prove the existence of scalar noise schedules βtβ_t that approach the CondOT noise schedule 1−t1-t at rate 1/S1/S and yield exact Gaussian sampling for every sufficiently large SS. Controlled Gaussian experiments illustrate the convergence rates and exact calibration.
Jul 13, 2026cs.LG

Velocity Scheduled Flow Matching

Flow matching trains a neural network to regress the conditional velocity along a linear interpolant between noise and data, and the number of network evaluations~(NFE) sets the cost of sampling. The straight-line interpolant carries an implicit choice: the sample moves at constant speed throughout the trajectory. We relax this choice and introduce Velocity Scheduled Flow Matching~(VSFM), which replaces the conditional target x1−x0x_1 - x_0 with v(t)(x1−x0)v(t)(x_1 - x_0) for any nonnegative profile v:[0,1]→R≥0v:[0,1]\to\mathbb{R}_{\geq 0} satisfying ∫01v dt=1\int_0^1 v\,dt = 1. We study six polynomial profiles drawn from motion planning. The first use of VSFM is at inference time: a pretrained linear flow-matching model can be sampled under any admissible profile by integrating its ODE on a non-uniform ττ-schedule, with no retraining and no additional computation; on CIFAR-10 this lowers FID by up to 19.8%19.8\%. Training from scratch under a braking profile gives a further reduction of 17.4%17.4\% at 44~NFE. Both gains follow from the local truncation error of the Euler integrator on the induced grid.