One-step generative models construct a static generator through iterative training-time transport. Existing transport objectives primarily assess distributional motion, although a neural generator needs to realize the requested sample displacements jointly through shared parameter updates. The training-time construction raises the question: \emph{once training becomes the iterative process that constructs the final one-step map, what to optimize: the next distributional move, or the route by which the finite generator learns the final map?} To address the question, we introduce \textbf{T}raining \textbf{D}ynamics \textbf{A}ction (\textbf{TDAction}), which selects transport targets according to local shared-parameter realization cost while retaining a prescribed level of distributional progress. We formulate the cost as a soft-terminal control problem and derive a closed-form Batch Tangent Action-to-Go value that accounts for parameter effort and terminal mismatch. The criterion captures cross-sample interactions omitted by independent pairwise costs; under isotropic mobility, the criterion agrees with quadratic Euclidean assignment for deterministic balanced couplings. Randomized tangent probes provide a low-rank implementation that constructs shared detached targets without adding an inference-time trajectory. Controlled studies examine the relationship between generator geometry, transport selection, and realized local action. On ImageNet 256×256, TDAction attains an FID below 1.1 without distillation.
Figures & tables
Figure 1 : Training Dynamics Action (TDAction). Illustrative two-dimensional transport comparison. (a) Distribution trajectories during training. (b) TDAction lowers target mismatch more rapidly through progress-preserving, generator-compatible displacements. (c) At matched progress, TDAction uses the least normalized action. All methods retain one-step inference.
Figure 2 : Predictive Action-to-Go for persistent one-step generation. (a) Local transports are absorbed during training to construct a persistent generator map. TDAction forecasts local generator mobilities, aggregates the mobilities into an Action-to-Go objective, and selects a predictive update (purple); the dashed gray path is a myopic local-update contrast. The resulting generator is still evaluated once at inference. (b) A conceptual toy action landscape illustrates why a locally attractive update need not minimize the remaining generator action: the predictive path uses the future-mobility estimate to target the terminal basin. The lower insets depict q(0) , q(k) , and pdata . Panel (b) is illustrative and does not report an empirical landscape or result.
Training Principle
Defining Object
Training-Time Rule
Drifting
Vdrift(q,p)
x+=x+ηVdrift(x)
Sinkhorn-Drifting
Sinkhorn-corrected field
Vε=Tq,pε−Tq,qε
W-Flow
Fε(q)=Sε(q,p)
V=−∇xδqδFε
TDAction (ours)
Vρ(π)=21dπ⊤(Kθ+ρI)−1dπ
πk⋆∈argπminVρ(π)
Table 1 : Comparison of training-time transport principles. The first three methods prescribe or derive distributional motion; TDAction selects how sufficient distributional progress is realized by the finite generator.
Figure 4
Figure 3 : Action-aware guidance under a finite generator budget. Dashed contours denote the KL-CFG reference target pw(x∣c) . The Reference row is the oracle action reference: under the same W-Flow guided velocity, finite generator geometry, and action budget, the reference row uses Kreal to compute the optimal persistent endpoint assignment. TDAction instead uses K with a 2∘ directional error and a 1% scale error; Drifting, Sinkhorn-Drifting, and W-Flow remain distribution-first dynamics. Right: five-seed mean W22 to the oracle action reference (top), to the reference CFG target (middle), and the mode-shape error (bottom); error bars show s.d.
Figure 4 : Laplacian-kernel toy transports. 8-Gaussians and Checkerboard across τ∈{0.01,0.05,0.1} for Drifting, Sinkhorn-Drifting, W-Flow, and TDAction. Left: final generated samples (tan) and target samples (blue). Right: W22 trajectories across five random seeds.
Figure 5 : Finite-generator trajectories in two toy transports. (a) Four methods under matched source/target distributions, held-out learner-response schedule, horizon, and action budget; Cross-Swap contrasts ambient-short and generator-easy maps, while Ring-Cycle repeats the contrast over eight modes. TDAction uses predictive Action-to-Go geometry with realized response held out from the mobility forecast. (b) Final empirical W22 across five seeds.
Probes
Frob. ↓
Angle ↓
Inverse ↓
Regret ↓
16
.484
3.96∘
.284
3.55×10−3
32
.314
2.23∘
.171
1.62×10−3
64
.207
1.52∘
.107
8.06×10−4
128
.142
1.07∘
.078
3.55×10−4
256
.103
0.75∘
.053
1.71×10−4
512
.072
0.53∘
.033
7.77×10−5
Table 2: Active-probe mobility identification (20 seeds). Frob. is relative Frobenius error, angle is the mean principal angle of the six-dimensional response subspace, inverse is relative operator error, and regret is the relative true-metric action-objective gap.
#Params
NFE
FID ↓
IS ↑
Multi-step Diffusion/Flows
ADM-G [ 8 ]
554M
250 × 2
4.59
186.7
DiT-XL/2 [ 32 ]
675M+49M
250 × 2
2.27
278.2
SiT-XL/2 [ 29 ]
675M+49M
250 × 2
2.06
270.3
SiT-XL/2+REPA [ 43 ]
675M+49M
250 × 2
1.42
305.7
LightningDiT-XL/2 [ 42 ]
675M+70M
250 × 2
1.35
295.3
Table 3 : Selected class-conditional ImageNet 256 × 256 context and TDAction model entries. Values and evaluation configurations in citation-based rows are reproduced from the cited sources and are not necessarily protocol-matched. TDAction rows identify the compared architectures. “#Params” denotes the number of parameters of “generator + decoder” when reported.
Method
Space
#Params
NFE
CFG
Eval. samples
FID ↓
IS ↑
Drifting Model, B/2
SD-VAE latent
133M+49M
1
1.10
50K
1.75
263.2
Drifting Model, L/2
SD-VAE latent
463M+49M
1
1.00
50K
1.54
258.9
W-Flow, B/2
SD-VAE latent
133M+49M
1
1.19
50K
1.52
271.8
W-Flow, L/2
SD-VAE latent
463M+49M
1
1.14
50K
1.35
272.5
W-Flow, XL/2
SD-VAE latent
679M+49M
1
1.09
50K
1.29
265.4
TDAction, B/2
SD-VAE latent
133M+49M
1
grid → 1.25
50K held-out
2.4637 †
251.2
Table 4 : Published baselines and a fixed-checkpoint TDAction diagnostic.
Figure 6 : Paired, uncurated samples from the frozen 56K-step EMA checkpoint with seed 2027 and 16 fixed ImageNet classes. Only the code-guidance scale changes.
Step
Updates
κ(Λ)
κ(Aˉ−1)
Angle
Plan churn
20K
156
1.12
1.259
–
–
40K
312
1.63
1.282
87.0∘
6.70×10−4
54K
421
1.17
1.262
88.1∘
3.47×10−4
56K
437
1.32
1.273
87.9∘
8.52×10−2
Table 5: Passive mobility audit for ImageNet-256 B/2 checkpoints. Plan churn is averaged over 20 fixed canonical-feature probe sets, and angle is the mean principal angle to the preceding checkpoint.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Principle
Functional
Induced rule
Distribution-energy-driven dynamics
Squared MMD
21MMD2(qt,p)
∫∇xk(x,y)p(dy)−∫∇xk(x,y)qt(dy)
KL divergence
DKL(qt∥p)
∇xlogp(x)−∇xlogqt(x)
Sinkhorn divergence
Sε(qt,p)
Tqt,pε(x)−Tqt,qtε(x)
Generator-action-driven transport selection
TDAction (ours)
Vρ(π)=21dπ⊤(Kθ+ρI)−1dπ
πk⋆∈argπminVρ(π),Pk(π)≥γPk(πE)
Appendix
Table 6 : Distributional-energy dynamics versus generator-action selection.
Drifting
W-Flow
Setting
B/2
L/2
B/2
L/2
XL/2
Generator and latent representation
architecture
DiT-B/2
DiT-L/2
DiT-B/2
DiT-L/2
DiT-XL/2
latent / patch
322×4 / 22
322×4 / 22
322×4 / 22
322×4 / 22
322×4 / 22
hidden dimension / depth
768 / 12
1024 / 24
768 / 12
1024 / 24
1152 / 28
register / style tokens
16 / 32
16 / 32
16 / 32
16 / 32
16 / 32
Appendix
Table 7 : Published ImageNet-256 configurations reported by Drifting [ 7 ] and W-Flow [ 16 ] . The listed values are reference-baseline values, not TDAction hyperparameters. Drifting reports CFG in α , whereas W-Flow reports w .
TDAction setting
B/2
L/2
XL/2
Canonical support and mobility estimation
canonical generated / real support
--
--
--
canonical encoder and normalization
--
--
--
control metric Rθ
--
--
--
primary / shadow probe counts
--
--
--
sketch rank r / floor σ2
--
--
--
Appendix
Table 8 : TDAction-specific planning record. All -- entries deliberately mark quantities not yet recorded by an audited TDAction ImageNet run; the entries are not values inherited from the published reference configurations.
Metric
W-Flow
TDAction
Relative change
2D Fréchet proxy ↓
0.04687±0.02334
0.04173±0.01888
−10.98%
Sliced Wasserstein ↓
0.34059±0.01483
0.33509±0.01264
−1.62%
Optimizer-time action ↓
70.1898±6.4149
62.3337±5.7703
−11.19%
Mode coverage ↑
8/8
8/8
–
Appendix
Table 9: Controlled 2D local-action study for TDAction. The endpoint values are not ImageNet FID.
Method
Fréchet proxy ↓
Action ↓
Action ratio ↓
W-Flow
0.12467
17.6794
1.0000
TDAction
0.13423
27.8792
1.7986
TDAction
0.11718
16.9900
0.8666
TDAction
0.11682
16.9607
0.8660
Appendix
Table 10: Three-seed 2D horizon screen for TDAction. The experiment concerns alternative planning settings and is not an ablation of the main TDAction implementation.
Setting
Code s
Actual w
FID ↓
IS ↑
Intrinsic
1.00
0.00
3.5589
197.76±4.94
Selected CFG
1.25
0.25
2.4637
251.15±4.88
Appendix
Table 11: Held-out 50K evaluation of the frozen 56K-step EMA checkpoint. Both rows use class-balanced sampling, seed 2027, and the same ImageNet-256 statistics; s=1.25 was selected only on the disjoint 10K/seed-0 sweep.
Figure 7 : Guidance trade-off at the frozen 56K-step EMA checkpoint. Error bars show IS split s.d., not training-seed uncertainty; code guidance satisfies s=w+1 .
We introduce a new framework for one-step generative modelling on finite state spaces. To extend drifting beyond continuous domains, we use discrete Wasserstein geometry to define a target-relative KL gradient flow over the transitions of a reversible Markov kernel. We realize this probability flow at the particle level through Markov jumps and amortize the resulting transport updates into a latent-conditioned generator, so that the iterative dynamics are required only during training while inference remains one-step. In a controlled setting where the underlying distributions and transport dynamics can be computed exactly, we verify KL dissipation, consistency between the particle dynamics and the probability flow, and the predicted numerical scaling. We further show that a finite-capacity neural generator can track these exact transport targets while retaining one-step generation. These results validate the basic construction and provide a foundation for scaling Discrete Drifting to structured discrete data.
Alessandro Micheli, Andrea Zerio, Samir Bhatt
Imperial College London London, United Kingdom · Department of Computer Science, Aalborg University, Copenhagen, Denmark; Centre for Frontier AI Research (CFAR), Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore · University of Copenhagen Copenhagen, Denmark
We propose a new framework for generative modeling based on a discrete-time stochastic control formulation of measure transport. Adapting classic results from control theory, we formulate our problem as a linear program whose dual variables correspond to the \emph{optimal value function} of the control problem, which directly encodes the optimal control policy. Exploiting this LP formulation, we develop an efficient simulation-free primal-dual algorithm for computing approximately optimal value functions and the associated \emph{value-driven transport} (VDT) policies which approximate the true optimal policy. We show that well-trained VDT policies enjoy numerous favorable properties in comparison with other state-of-the-art methods based on flows, diffusions, or Schrödinger bridges: they lead to straight transport paths which can be simulated quickly and robustly, and can be enhanced in all the same ways as diffusion and flow-based models (e.g., conditional generation, classifier-free guidance, unpaired data-to-data translation are all easy to incorporate). We evaluate our methodology in a range of experiments, with results that indicate strong performance and good potential for scalability.
Pablo Moreno-Muñoz, Adrian Müller, Gergely Neu
Universitat Pompeu Fabra · Barcelona, Spain · ETH Zürich +2
Diffusion models and flow-based methods have shown impressive generative capability, especially for images, but their sampling is expensive because it requires many iterative updates. We introduce W-Flow, a framework for training a generator that transforms samples from a simple reference distribution into samples from a target data distribution in a single step. This is achieved in two steps: we first define an evolution from the reference distribution to the target distribution through a Wasserstein gradient flow that minimizes an energy functional; second, we train a static neural generator to compress this evolution into one-step generation. We instantiate the energy functional with the Sinkhorn divergence, which yields an efficient optimal-transport-based update rule that captures global distributional discrepancy and improves coverage of the target distribution. We further prove that the finite-sample training dynamics converge to the continuous-time distributional dynamics under suitable assumptions. Empirically, W-Flow sets a new state of the art for one-step ImageNet 256×256 generation, achieving 1.29 FID, with improved mode coverage and domain transfer. Compared to multi-step diffusion models with similar FID scores, our method yields approximately 100× faster sampling. These results show that Wasserstein gradient flows provide a principled and effective foundation for fast and high-fidelity generative modeling.