Beyond Gradient Flow: Identifiability and Recovery from Distribution Snapshots
Authors: Nam D. Nguyen, Valeriya Malysheva
Organizations: VIB, Center for Molecular Neurology, Antwerp, Belgium · VIB, Center for AI and Computational Biology, Leuven, Belgium · Faculty of Pharmaceutical, Biomedical and Veterinary Sciences, University of Antwerp, Antwerp, Belgium · Research Foundation – Flanders (FWO), Brussels, Belgium · Trinity Hall, University of Cambridge, Cambridge, UK
Inferring dynamics from snapshots of evolving distributions is fundamentally underdetermined: the Fokker-Planck equation constrains the drift F only through its score-weighted divergence ∇⋅F+F⋅∇logρ, leaving a ρ-solenoidal gauge invisible to any single-time constraint. Time-indexed transport formulations cannot resolve this ambiguity: every admissible marginal path admits a curl-free explanation, minimum-action reconstruction selects it, and marginal fit alone cannot distinguish dynamically inequivalent explanations. Requiring one autonomous field to explain several marginals instead makes part of the hidden circulation visible as ∇logρ changes across marginals. Separating instantaneous Fokker-Planck source constraints from the snapshot experiment, we show that the source constraints identify the field modulo the kernel of a stacked score-weighted divergence operator. For generic Gaussian shape variation, source constraints at K≥m time points in intrinsic dimension m eliminate every polynomial gauge direction, whereas finitely many density snapshots alone admit aliasing; we give the obstruction explicitly. At a Gaussian anchor, for Sobolev smoothness s and n samples per time point, we derive a conditional lower rate (nK)−2s/(2s+m+1) for the tangent snapshot experiment, with a matching upper rate in a degreewise benchmark. Strong-form fitting is non-orthogonal to score error and cannot be repaired by spectral filtering. Instead, we estimate using smooth test functions while retaining the known diffusion term, and derive a finite-sample bound that separates sampling error from fixed-grid quadrature bias. Planted-circulation experiments confirm the predicted gauge contraction and expose a design tension between cross-slice information and covariance-aware whitening.
Figures & tables
Figure 1: Identical marginal flow, different dynamics. (a) A flow with a planted rotational component and a curl-free transport generate the same Gaussian marginals. (b) On the reference model of Tab. 1 , linearized-GLS decays at the n−1/2 rate whereas displacement interpolation has solenoidal relative error 1 for every n , including n=∞ . Both match the observed marginals to machine precision, so marginal fit cannot distinguish them.
estimator
ingredients
n=40 k
n=200 k
n=40 k, 12 seeds
gradient-only / displacement
—
1.000
1.000
—
weak, unweighted
moments
—
—
0.734±0.157
weak, diagonal weights
moments
0.893
0.303
0.782±0.199
weak, GLS
moments
0.651
0.282
0.751±0.260
parametric reference
exact linear model
0.511
0.179
0.308±0.088
strong form, oracle nuisances
exact score & source
—
—
0.535±0.177
Table 1: Planted-curl OU benchmark ( d=12 , m=4 , K=5 , σ2=ν=1 ). Entries are \relerr . Only weak rows are nuisance-free: the parametric reference is given the symmetric–antisymmetric tie, while strong-form rows use the exact score and temporal source. Multi-seed entries are mean ± s.d.; the four snapshot-only arms share identical draws over 12 paired seeds, whereas strong-form rows use a separate 8 -seed set. See App. Y for single-seed and seed-spread caveats and solver variants.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
claim
status
scope and what it rests on
Prop. 1
P
Poincaré inequality, ∂tlogρt∈L2(ρt) ; App. C
Cor. 2
P
(kerAρ)⊥={∇ϕ} ; single time
Cor. 18
P
Gaussian marginals; closed-form Benamou–Brenier
Cor. 19
P
population level; (i) Gaussian slices, exact OT coupling, noiseless interpolant; (ii) any (possibly non-gradient) reference drift, smooth positive Schrödinger potentials
Prop. 4
P
source constraints only, at finitely many or a continuum of times; forward uniqueness for the continuum converse
Ex. 5
P
explicit; shows kerAK is not necessary for snapshot equivalence
Appendix
Table 2: Hypothesis ledger. Status of every numbered claim and what it rests on (App. A ).
parameterization
params
rank
cond.
error
reached
antisymmetric potential
5226
144
3.1×104
0.542±0.105
4/5
generator frame
144
144
2.7×104
0.521±0.092
5/5
+ gauge reduction
116
116
1.0×103
0.521±0.092
5/5
+ audit preconditioning
116
116
1.0
0.963±0.385
2/5
+ tangent-only support
84
84
9.2×102
0.520±0.081
5/5
Gauss–Newton (exact solve)
—
—
—
0.479±0.194
—
Appendix
Table 3: Curl parameterizations. Solenoidal relative error after 8000 descent steps; “reached” counts seeds whose loss came within 2.7% of the Gauss–Newton value. Conditioning is of the whitened design in that parameterization.
Functional flow matching is posed on distributions of functions but implemented from finitely many coefficients or point values. Under scattered or adaptive refinement, the resulting conditioning sigma-algebras need not be nested, so martingale convergence does not justify the sensor limit. We prove strong L2 convergence of finite conditional velocity targets for every strongly consistent sequence of finite-rank reconstructions, with quantitative bounds for orthogonal projections and a point-sensor extension through a regularity space. For learned flows, coupling directly to a population superposition path yields an end-to-end Wasserstein bound without assuming uniqueness of the population finite-dimensional ODE. We verify sensor-independent constants for a normalized quadrature neural operator, including globally Lipschitz activations through an explicit magnitude recurrence. A noncommuting trace-class Gaussian example gives boundary multiplier 0 under projected restriction and 0.72 under exact conditioning. A spatial regularity--cubature certificate closes the operator-realization term, a Bernstein argument gives a O(n−1) excess-risk term for fixed model dimension and envelopes, and an exactly realizable clipped Gaussian scaling specialization yields an explicit end-to-end rate.
Lennon J. Shikhman
School of Computer Science, College of Computing Georgia Institute of Technology Atlanta, Georgia 30332, USA
Flow matching generates samples by gradually transforming noise into data. In practice, using a finite number of sampling steps introduces a numerical error that depends on the chosen schedule. We study this dependence for Gaussian targets and the explicit midpoint sampling method, using the exact flow field. We measure sampling error by the squared Wasserstein distance between the target distribution and the final distribution produced by the midpoint sampler. We show that the standard conditional optimal transport (CondOT) schedule cancels the leading midpoint error and improves the general convergence bound, even when the sampling steps are unequally spaced. On a uniform grid of S sampling steps, we fix the signal schedule at αt=t and prove the existence of scalar noise schedules βt that approach the CondOT noise schedule 1−t at rate 1/S and yield exact Gaussian sampling for every sufficiently large S. Controlled Gaussian experiments illustrate the convergence rates and exact calibration.
Reconstructing PDE-governed fields from sparse and irregular measurements is challenging due to their ill-posed nature. Deterministic surrogates are trained on dense fields that struggle with limited measurements and uncertainty quantification. Generative models, by learning distributions over spatiotemporal fields, can better handle sparsity and uncertainty. However, existing generative approaches enforce data consistency and PDE constraints simultaneously via sampling-time gradient guidance, resulting in slow and unstable inference. To this end, we propose PerFlow, a Physics-embedded rectified Flow for efficient sparse reconstruction and uncertainty quantification of spatiotemporal dynamics. PerFlow decouples observation conditioning from physics enforcement, performing guidance-free conditioning by feeding observations into rectified-flow dynamics while embedding hard physics via a constraint-preserving projection (e.g., incompressibility or conservation). Theoretically, we establish invariance guarantees to ensure that trajectories remain on the physics-consistent manifold throughout sampling. Experiments on various PDE systems demonstrate competitive reconstruction accuracy with sound physics consistency, while enabling efficient conditional sampling (e.g., 50 steps) and up to 320x faster inference than 2000-step guided diffusion baselines.
Hao Zhou, Rui Zhang, Han Wan +1
Gaoling School of Artificial Intelligence, Renmin University of China