Learning population dynamics from unpaired temporal marginals is an ill-posed inverse problem that requires structural assumptions on the underlying dynamics. We introduce Discrete Action Matching (DAM), a finite-state counterpart of Action Matching based on discrete Wasserstein geometry. For a prescribed marginal path and transport geometry, we derive an action-minimization objective for its canonical minimum-kinetic-energy current. Our key observation is that the density dependence of the discrete action reduces to neighboring density ratios. Along an empirical interpolation of the snapshots, DAM first estimates these ratios and then learns an action potential. The learned fields also define a graph-supported Markov sampler. Experiments on controlled synthetic dynamics and real mouse gastrulation data evaluate marginal reconstruction and interpolation. Additional experiments approximate numerical surface-transport paths from paired samples.
Figures & tables
Figure 2: Synthetic reconstruction across eleven graph families with N=60 , M=1,000 samples per split and time, and β=0.1 . DAM uses logarithmic mobility; lower is better.
Figure 3: Gastrulation interpolation against omitted smoothed atlas marginals. DAM uses harmonic mobility; lower is better.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Role
Continuous sample space
Finite graph
Probability state
q∈P+(Ω)
p∈P+(I)
Tangent vector
σ∈TqP+(Ω),∫Ωσdx=0
σ∈TpP+(I),1⊤σ=0
Potential
[s]∈C∞(Ω)/R
[S]∈RN/span{1}
Gradient
∇s
(∇GS)ij=ωij(Si−Sj)
Potential-to-tangent map
−div(q∇s)
−divG(θ(p)∇GS)=Lθ(p)S
Continuity equation
q˙t=−div(qt∇st)
p˙t=Lθ(pt)St
Appendix
Table 1: Continuous–discrete dictionary for the dynamical transport structures underlying Action Matching.
Setting
DAM
FEGF
Architecture
16-dimensional node embedding, scalar time, two width-64 ReLU layers, scalar output
Synthetic: edge network with embedding 16 and width 64; gastrulation: time-indexed scalar table
Learning rate
5⋅10−4
Synthetic: 5⋅10−4 ; gastrulation: 10−2
Batch size
128
64
Maximum epochs
Ratio: 60; action: 100
100
Selection
Ratio validation loss; action observed-time validation TV
Observed-time validation TV
Early-stopping patience
Ratio: 10 epochs; action: none
15 epochs
Appendix
Table 2: Synthetic and gastrulation training settings. All models use Adam, float64 arithmetic, weight decay 10−6 , gradient-clipping norm 1 and 200 updates per epoch.
Figure 4: Synthetic sweeps with logarithmic mobility. Bars average Hellinger errors over eleven graph families; whiskers show sample SD across three seeds. Each panel varies one factor from the base setting. Smaller sample budgets share draws.
Dataset
Mobility
H (L)
H (Hst)
TV (L)
TV (Hst)
Synthetic
Arithmetic
0.110
0.110
0.106
0.107
Geometric
0.109
0.108
0.106
0.106
Harmonic
0.109
0.108
0.106
0.106
Logarithmic
0.109
0.109
0.106
0.106
Gastrulation
Arithmetic
0.285
0.262
0.268
0.242
Geometric
0.234
0.229
0.220
0.215
Appendix
Table 3: Mean errors for separately fitted mobilities. Synthetic results average eleven base-setting graphs and three seeds; gastrulation results average seven omitted stages and three seeds. L: learned ratios; Hst: histograms. Lower is better; values are rounded to three decimals.
Figure 5: Hellinger errors for learned-ratio and histogram DAM, refitted for each mobility. Bars and whiskers show means and sample SD across three seeds after averaging graphs or omitted stages. Panels use separate vertical scales.
Figure 6: Gastrulation TV errors against omitted, smoothed atlas marginals. Each fold uses eight retained stages, including later observations. DAM uses harmonic mobility; FEGF uses its stationary-relative logarithmic geometry. Whiskers show sample SD across three sampling/optimization seeds, rather than biological replicates.
Method
Sphere
Hand
DAM, histogram
0.169
0.246
DAM, local + global ratios
0.192
0.250
Tuned FEGF
0.564
0.843
Appendix
Table 4: Mean TV at 16 omitted times against independent test-particle marginals, seed 42; lower is better.
Figure 7: Sphere transport, seed 42. Rows: numerical teacher, tuned FEGF, DAM with local + global ratios, and histogram DAM. Columns: t=0,0.25,0.5,0.75,1 . Colors show area-normalized probability density.
Figure 8: Hand transport, seed 42, using the row order, display times and density-rendering convention of Figure 7 .
Figure 9: Accuracy versus CPU cost on sphere (top) and hand (bottom): training (left) and marginal inference (right, logarithmic time axis). TV averages 32 noninitial times. Whiskers span three runtime repetitions around median times. Tuned FEGF is inference-only.