Learning population dynamics from unpaired temporal marginals is an ill-posed inverse problem that requires structural assumptions on the underlying dynamics. We introduce Discrete Action Matching (DAM), a finite-state counterpart of Action Matching based on discrete Wasserstein geometry. For a prescribed marginal path and transport geometry, we derive an action-minimization objective for its canonical minimum-kinetic-energy current. Our key observation is that the density dependence of the discrete action reduces to neighboring density ratios. Along an empirical interpolation of the snapshots, DAM first estimates these ratios and then learns an action potential. The learned fields also define a graph-supported Markov sampler. Experiments on controlled synthetic dynamics and real mouse gastrulation data evaluate marginal reconstruction and interpolation. Additional experiments approximate numerical surface-transport paths from paired samples.
Figures & tables
Figure 2: Synthetic reconstruction across eleven graph families with N=60 , M=1,000 samples per split and time, and β=0.1 . DAM uses logarithmic mobility; lower is better.
Figure 3: Gastrulation interpolation against omitted smoothed atlas marginals. DAM uses harmonic mobility; lower is better.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Role
Continuous sample space
Finite graph
Probability state
q∈P+(Ω)
p∈P+(I)
Tangent vector
σ∈TqP+(Ω),∫Ωσdx=0
σ∈TpP+(I),1⊤σ=0
Potential
[s]∈C∞(Ω)/R
[S]∈RN/span{1}
Gradient
∇s
(∇GS)ij=ωij(Si−Sj)
Potential-to-tangent map
−div(q∇s)
−divG(θ(p)∇GS)=Lθ(p)S
Continuity equation
q˙t=−div(qt∇st)
p˙t=Lθ(pt)St
Appendix
Table 1: Continuous–discrete dictionary for the dynamical transport structures underlying Action Matching.
Setting
DAM
FEGF
Architecture
16-dimensional node embedding, scalar time, two width-64 ReLU layers, scalar output
Synthetic: edge network with embedding 16 and width 64; gastrulation: time-indexed scalar table
Learning rate
5⋅10−4
Synthetic: 5⋅10−4 ; gastrulation: 10−2
Batch size
128
64
Maximum epochs
Ratio: 60; action: 100
100
Selection
Ratio validation loss; action observed-time validation TV
Observed-time validation TV
Early-stopping patience
Ratio: 10 epochs; action: none
15 epochs
Appendix
Table 2: Synthetic and gastrulation training settings. All models use Adam, float64 arithmetic, weight decay 10−6 , gradient-clipping norm 1 and 200 updates per epoch.
Figure 4: Synthetic sweeps with logarithmic mobility. Bars average Hellinger errors over eleven graph families; whiskers show sample SD across three seeds. Each panel varies one factor from the base setting. Smaller sample budgets share draws.
Dataset
Mobility
H (L)
H (Hst)
TV (L)
TV (Hst)
Synthetic
Arithmetic
0.110
0.110
0.106
0.107
Geometric
0.109
0.108
0.106
0.106
Harmonic
0.109
0.108
0.106
0.106
Logarithmic
0.109
0.109
0.106
0.106
Gastrulation
Arithmetic
0.285
0.262
0.268
0.242
Geometric
0.234
0.229
0.220
0.215
Appendix
Table 3: Mean errors for separately fitted mobilities. Synthetic results average eleven base-setting graphs and three seeds; gastrulation results average seven omitted stages and three seeds. L: learned ratios; Hst: histograms. Lower is better; values are rounded to three decimals.
Figure 5: Hellinger errors for learned-ratio and histogram DAM, refitted for each mobility. Bars and whiskers show means and sample SD across three seeds after averaging graphs or omitted stages. Panels use separate vertical scales.
Figure 6: Gastrulation TV errors against omitted, smoothed atlas marginals. Each fold uses eight retained stages, including later observations. DAM uses harmonic mobility; FEGF uses its stationary-relative logarithmic geometry. Whiskers show sample SD across three sampling/optimization seeds, rather than biological replicates.
Method
Sphere
Hand
DAM, histogram
0.169
0.246
DAM, local + global ratios
0.192
0.250
Tuned FEGF
0.564
0.843
Appendix
Table 4: Mean TV at 16 omitted times against independent test-particle marginals, seed 42; lower is better.
Figure 7: Sphere transport, seed 42. Rows: numerical teacher, tuned FEGF, DAM with local + global ratios, and histogram DAM. Columns: t=0,0.25,0.5,0.75,1 . Colors show area-normalized probability density.
Figure 8: Hand transport, seed 42, using the row order, display times and density-rendering convention of Figure 7 .
Figure 9: Accuracy versus CPU cost on sphere (top) and hand (bottom): training (left) and marginal inference (right, logarithmic time axis). TV averages 32 noninitial times. Whiskers span three runtime repetitions around median times. Tuned FEGF is inference-only.
Flow matching trains a neural velocity field by regression against a target velocity associated with a prescribed probability path connecting a simple initial distribution to the data distribution. A central design choice is the path itself. Existing constructions, including rectified and optimal-transport-based paths, transport samples along straight lines between coupled endpoints and thus cover only a narrow class of dynamics. We observe that this corresponds to the simplest case of the least-action principle in classical mechanics, in which the kinetic Lagrangian yields free-particle straight-line trajectories. Building on this observation, we propose Lagrangian flow matching, a physics-based framework in which the probability path and velocity field are determined by minimizing the action of a general Lagrangian subject to the continuity equation and the prescribed endpoints. We show that this dynamic problem admits an equivalent static optimal transport (OT) formulation, yielding a family of simulation-free training objectives that recover OT-based flow matching as the kinetic special case and the trigonometric variance-preserving diffusion path as the harmonic-oscillator case. More general Lagrangians give rise to new probability paths and velocity fields, and numerical experiments show that they induce meaningful changes in the learned dynamics while remaining competitive with existing conditional flow matching models.
Shukai Du, Junzhe Zhang, Yiming Li
Department of Mathematics, Syracuse University · Department of EECS, Syracuse University
The population dynamics of molecules, cells, and organisms are governed by a number of unknown forces. In the last decade, population dynamics have predominantly been modeled with Wasserstein gradient flows. However, since gradient flows minimize free energy, they fail to capture important dynamical properties, such as periodicity. In this work, we propose a change in perspective by considering dynamics that minimize a population-level action under a damped Wasserstein Lagrangian. By deriving the corresponding Hamiltonian equations of motion, we formalize Wasserstein Lagrangian Mechanics, a structured class of second-order dynamics that encompasses classical mechanics, quantum mechanics, and gradient flows. We then propose WLM as the first algorithm that learns these second-order dynamics from observed marginals, without specifying the Lagrangian. By directly learning the population mechanics, WLM can both forecast and interpolate unseen marginals, and outperforms existing gradient flow and flow matching methods across a wide range of dynamics, including vortex dynamics, embryonic development, and flocking.
Vincent Guan, Lazar Atanackovic, Kirill Neklyudov
University of British Columbia · Mila - Quebec AI Institute · University of Alberta +4
Learning the dynamics of a process given sampled observations at several time points is an important but difficult task in many scientific applications. When no ground-truth trajectories are available, but one has only snapshots of data taken at discrete time steps, the problem of modelling the dynamics, and thus inferring the underlying trajectories, can be solved by multi-marginal generalisations of flow matching algorithms. This paper proposes a novel flow matching method that overcomes the limitations of existing multi-marginal trajectory inference algorithms. Our proposed method, ALI-CFM, uses a GAN-inspired adversarial loss to fit neurally parametrised interpolant curves between source and target points such that the marginal distributions at intermediate time points are close to the observed distributions. The resulting interpolants are smooth trajectories that, as we show, are unique under mild assumptions. These interpolants are subsequently marginalised by a flow matching algorithm, yielding a trained vector field for the underlying dynamics. We showcase the versatility and scalability of our method by outperforming the existing baselines on spatial transcriptomics and cell tracking datasets, while performing on par with them on single-cell trajectory prediction. Code: https://github.com/mmacosha/adversarially-learned-interpolants.
Oskar Kviman, Kirill Tamogashev, Nicola Branchini +3