One-Step Generative Modeling via Training Dynamics Action
Organizations: Department of Artificial Intelligence, Westlake University · Hangzhou Normal University
Abstract
One-step generative models construct a static generator through iterative training-time transport. Existing transport objectives primarily assess distributional motion, although a neural generator needs to realize the requested sample displacements jointly through shared parameter updates. The training-time construction raises the question: \emph{once training becomes the iterative process that constructs the final one-step map, what to optimize: the next distributional move, or the route by which the finite generator learns the final map?} To address the question, we introduce \textbf{T}raining \textbf{D}ynamics \textbf{A}ction (\textbf{TDAction}), which selects transport targets according to local shared-parameter realization cost while retaining a prescribed level of distributional progress. We formulate the cost as a soft-terminal control problem and derive a closed-form Batch Tangent Action-to-Go value that accounts for parameter effort and terminal mismatch. The criterion captures cross-sample interactions omitted by independent pairwise costs; under isotropic mobility, the criterion agrees with quadratic Euclidean assignment for deterministic balanced couplings. Randomized tangent probes provide a low-rank implementation that constructs shared detached targets without adding an inference-time trajectory. Controlled studies examine the relationship between generator geometry, transport selection, and realized local action. On ImageNet , TDAction attains an FID below without distillation.
Figures & tables
| Training Principle | Defining Object | Training-Time Rule |
|---|---|---|
| Drifting | ||
| Sinkhorn-Drifting | Sinkhorn-corrected field | |
| W-Flow | ||
| TDAction (ours) |
| Probes | Frob. | Angle | Inverse | Regret |
|---|---|---|---|---|
| 16 | .484 | .284 | ||
| 32 | .314 | .171 | ||
| 64 | .207 | .107 | ||
| 128 | .142 | .078 | ||
| 256 | .103 | .053 | ||
| 512 | .072 | .033 |
| #Params | NFE | FID | IS | |
| Multi-step Diffusion/Flows | ||||
| ADM-G [ 8 ] | 554M | 250 2 | 4.59 | 186.7 |
| DiT-XL/2 [ 32 ] | 675M+49M | 250 2 | 2.27 | 278.2 |
| SiT-XL/2 [ 29 ] | 675M+49M | 250 2 | 2.06 | 270.3 |
| SiT-XL/2+REPA [ 43 ] | 675M+49M | 250 2 | 1.42 | 305.7 |
| LightningDiT-XL/2 [ 42 ] | 675M+70M | 250 2 | 1.35 | 295.3 |
| Method | Space | #Params | NFE | CFG | Eval. samples | FID | IS |
|---|---|---|---|---|---|---|---|
| Drifting Model, B/2 | SD-VAE latent | 133M+49M | 1 | 1.10 | 50K | 1.75 | 263.2 |
| Drifting Model, L/2 | SD-VAE latent | 463M+49M | 1 | 1.00 | 50K | 1.54 | 258.9 |
| W-Flow, B/2 | SD-VAE latent | 133M+49M | 1 | 1.19 | 50K | 1.52 | 271.8 |
| W-Flow, L/2 | SD-VAE latent | 463M+49M | 1 | 1.14 | 50K | 1.35 | 272.5 |
| W-Flow, XL/2 | SD-VAE latent | 679M+49M | 1 | 1.09 | 50K | 1.29 | 265.4 |
| TDAction, B/2 | SD-VAE latent | 133M+49M | 1 | grid 1.25 | 50K held-out | 2.4637 † | 251.2 |
| Step | Updates | Angle | Plan churn | ||
|---|---|---|---|---|---|
| 20K | 156 | 1.12 | 1.259 | – | – |
| 40K | 312 | 1.63 | 1.282 | ||
| 54K | 421 | 1.17 | 1.262 | ||
| 56K | 437 | 1.32 | 1.273 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Principle | Functional | Induced rule |
|---|---|---|
| Distribution-energy-driven dynamics | ||
| Squared MMD | ||
| KL divergence | ||
| Sinkhorn divergence | ||
| Generator-action-driven transport selection | ||
| TDAction (ours) | ||
| Drifting | W-Flow | ||||
| Setting | B/2 | L/2 | B/2 | L/2 | XL/2 |
| Generator and latent representation | |||||
| architecture | DiT-B/2 | DiT-L/2 | DiT-B/2 | DiT-L/2 | DiT-XL/2 |
| latent / patch | / | / | / | / | / |
| hidden dimension / depth | 768 / 12 | 1024 / 24 | 768 / 12 | 1024 / 24 | 1152 / 28 |
| register / style tokens | 16 / 32 | 16 / 32 | 16 / 32 | 16 / 32 | 16 / 32 |
| TDAction setting | B/2 | L/2 | XL/2 |
| Canonical support and mobility estimation | |||
| canonical generated / real support | -- | -- | -- |
| canonical encoder and normalization | -- | -- | -- |
| control metric | -- | -- | -- |
| primary / shadow probe counts | -- | -- | -- |
| sketch rank / floor | -- | -- | -- |
| Metric | W-Flow | TDAction | Relative change |
|---|---|---|---|
| 2D Fréchet proxy | |||
| Sliced Wasserstein | |||
| Optimizer-time action | |||
| Mode coverage | – |
| Method | Fréchet proxy | Action | Action ratio |
|---|---|---|---|
| W-Flow | |||
| TDAction | |||
| TDAction | |||
| TDAction |
| Setting | Code | Actual | FID | IS |
|---|---|---|---|---|
| Intrinsic | 1.00 | 0.00 | 3.5589 | |
| Selected CFG | 1.25 | 0.25 | 2.4637 |