Parametric Trajectory Distillation for Few-Step Video Generation
Organizations: EPFL · NVIDIA · Stanford University
Abstract
Video diffusion and flow models require many sequential evaluations, making generation computationally expensive. Few-step distillation reduces this cost but poses a capacity allocation problem: a student must match the teacher's iterative generation with far less sequential computation. Existing trajectory methods ask the student to reproduce teacher transitions that are highly curved at high noise, which can exceed its capacity and degrade fine detail. We introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path. PTD is designed to let the learned curvature adapt to the backbone's predictive capacity, preserving motion and diversity. The curvature head is used only in training; inference keeps the original backbone architecture. On Wan2.1-14B, four-step PTD sets a new state of the art for trajectory distillation, significantly improving dynamic quality and naturalness over PDD, the best-performing trajectory-only method on this model, under the same training setting. On the 33B audio-video MiniMax-H3, LoRA-trained PTD significantly improves diversity and naturalness over the state-of-the-art LightX2V Turbo. Blinded human votes give PTD 55.1% and 63.4% preference shares against PDD and LightX2V Turbo. Project page: https://alan-lanfeng.github.io/PTD/.
Figures & tables
| Method | NFE | VBench | Diversity | |||
|---|---|---|---|---|---|---|
| Overall | Quality | Semantic | V-JEPA 2 | VideoMAE | ||
| Published reference results | ||||||
| UniPC ∗ (Teacher) | 83.90 | 84.56 | 81.24 | 0.1263 | 0.0250 | |
| rCM † | 4 | 84.92 | 85.43 | 82.88 | — | — |
| AnyFlow ∗ | 4 | 84.95 | 85.70 | 81.92 | 0.0786 | 0.0130 |
| DMD2 ∗∗ (FastGen) | 4 | 84.40 | 85.16 | 81.34 | 0.0568 | 0.0095 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Velocity target | Target query states | Student velocity | Deployed update | |
|---|---|---|---|---|
| LMD | Frozen teacher | Predicted map states | JVP of a two-time network | Two-time map |
| TVM | Own EMA velocity, plus flow matching on data | Predicted terminal states | JVP of a two-time network, backpropagated | Two-time map |
| -Flow | Frozen teacher | Detached policy rollout, numerically integrated | Closed-form policy | Policy integrated with substeps |
| PDD | Teacher midpoint RK target | Teacher solver stages | Discrete interval outputs | Fused interval outputs |
| PTD | Frozen teacher | Detached student curve, in closed form | Analytic basis derivative |
| Dimension | PTD | Baseline | 95% CI | W/T/L | |
| MiniMax-H3: PTD at 150 updates vs. released LightX2V Turbo v1.1 | |||||
| Dynamic quality | 2.70 | 2.61 | 45/27/28 | ||
| Static quality | 3.04 | 3.06 | 6/86/8 | ||
| Naturalness | 3.09 | 2.98 | 28/68/4 | ||
| Diversity | 2.81 | 1.55 | 93/6/1 | ||
| Wan2.1-14B: PTD at 64 updates vs. our PDD reproduction at 96 updates | |||||
| Prompt group | Dynamic quality | Static quality | Naturalness | Diversity | |
|---|---|---|---|---|---|
| MiniMax-H3 | |||||
| Live-action (V50) | 50 | ||||
| Action (D50) | 50 | ||||
| Wan2.1-14B | |||||
| Live-action (V50) | 50 | ||||
| Action (D50) | 50 | ||||
| Method | Updates | Overall | Quality | Semantic |
|---|---|---|---|---|
| PDD | 64 | 83.29 | 84.43 | 78.71 |
| PDD (best evaluated) | 96 | 83.67 | 85.24 | 77.40 |
| PDD | 256 | 82.83 | 83.21 | 81.29 |
| PTD (best evaluated) | 64 | 84.57 | 85.44 | 81.10 |
| Dimension | PDD @96 | PTD @64 | |
| Quality dimensions | |||
| Subject consistency | 94.45 | 95.79 | |
| Background consistency | 95.51 | 96.34 | |
| Temporal flickering | 98.74 | 98.80 | |
| Motion smoothness | 98.86 | 98.37 | |
| Dynamic degree | 85.83 | 76.11 | |
| Subset | PTD wins | Ties | Losses | Share | Non-tie win rate | 95% CI | ||
| MiniMax-H3: PTD at 150 updates vs. LightX2V Turbo v1.1 | ||||||||
| All | 82 | 45 | 14 | 23 | 63.4% | 66% | [54, 76] | 0.010 |
| Action (D50) | 43 | 22 | 10 | 11 | 62.8% | 67% | [50, 80] | 0.08 |
| Live-action (V50) | 39 | 23 | 4 | 12 | 64.1% | 66% | [49, 79] | 0.09 |
| Wan2.1-14B: PTD at 64 updates vs. PDD at 96 updates | ||||||||
| All | 79 | 33 | 21 | 25 | 55.1% | 57% | [44, 69] | 0.36 |
| Hop | Student MSE | Two queries | Three queries | ||
|---|---|---|---|---|---|
| MSE | Relative | MSE | Relative | ||
| 0 | 0.3346 | 0.04218 | 12.6% | 0.02224 | 6.6% |
| 1 | 0.0114 | 0.00155 | 13.6% | 0.00069 | 6.1% |
| 2 | 0.0036 | 0.00048 | 13.3% | 0.00021 | 5.8% |
| Hop 0 | Hop 1 | ||||||||
| Updates | Method | ||||||||
| 32 | PTD | 0.047 | 0.200 | 0.010 | 0.234 | 0.030 | 0.020 | 0.001 | 1.485 |
| PDD | 0.017 | 0.139 | 0.002 | 0.123 | 0.025 | 0.023 | 0.001 | 1.065 | |
| 64 | PTD | 0.081 | 0.171 | 0.014 | 0.476 | 0.046 | 0.028 | 0.001 | 1.668 |
| PDD | 0.024 | 0.139 | 0.003 | 0.172 | 0.025 | 0.023 | 0.001 | 1.090 | |
| 96 | PTD | 0.113 | 0.182 | 0.021 | 0.621 | 0.058 | 0.031 | 0.002 | 1.879 |
| Parameterization | Ridge | Adam | |||
|---|---|---|---|---|---|
| Polynomial coordinates, | |||||
| Legendre, natural scale | 8 | 7.9 | 45 | 50 | 46 |
| Legendre, unit RMS | 8 | 6.3 | 45 | 50 | 48 |
| Integrated Legendre, unit RMS | 8 | 1.05 | 45 | 49 | 48 |
| Bernstein (Bézier), unit RMS | 8 | 792 | 45 | 50 | 49 |
| Legendre, unit RMS | H3 probe (%) Adam | |
| 1 | 1.0 | 39 |
| 2 | 1.0 | 40–42 |
| 3 | 1.6 | 44–48 |
| 4 | 2.2 | 45–49 |
| 8 | 6.3 | 46–49 |