Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nyström--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.
Figures & tables
Figure 1 : Demonstration of Elastic Forcing. Instead of learning the distribution from teacher provided real and fake scores, Elastic Forcing directly learns autoregressive video rollouts from reference samples
Figure 2 : MMD in different representation spaces. Rows use DINOv3, VideoMAE, V-JEPA 2, and Wan VAE features, respectively; columns show sampled frames in temporal order. Wan VAE matching exhibits distortions in the helmet and clothing. The pretrained encoders preserve more coherent subjects, with different degrees and types of visible motion. This qualitative example complements the metric breakdown in Table 8 .
Figure 3 : A 2D toy experiment illustrating hybrid MMD estimation.
Figure 4 : Qualitative comparison with CausVid and Self Forcing. Red text highlights key prompt requirements, and red circles indicate representative failure cases. In these examples, our method demonstrates more faithful action execution, better preservation of subject counts, and more coherent interactions across frames.
Figure 5 : Scaling to 14B. Our streaming 14B model produces better results than our streaming 1.3B models in complex scenarios.
Figure 6Figure 7
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
1.3B model
14B model
Backbone
Wan2.1-T2V-1.3B
Wan2.1-T2V-14B
Initialization
Self-Forcing ODE checkpoint
Krea ODE checkpoint
Denoising steps
4 per chunk
5 per chunk
Post-training updates
150(three encoder)/100(two encoder)
80
Hardware
4 NVIDIA H200 GPUs
8 NVIDIA H200 GPUs
Appendix
Table 6 : Reported settings for standard text-to-video post-training. Both model scales use an ODE-initialized checkpoint.
Encoder
Tokens per video
Dimension per token
DINOv3
1024
1024
V-JEPA2
1024
1024
VideoMAE v2
1280
1024
Loss
(10NDINOv3+NV-JEPA2+NVideoMAE)/3
Nyström:MC hybrid
2:1
Appendix
Table 7 : Encoder settings for post-training.
Figure 8 : Effect of incorporating DINOv3 features. Adding DINOv3 supervision improves perceptual quality by preserving sharper details and more consistent object geometry. Without DINOv3, the generated content becomes overly smooth and may exhibit structural tearing around the horse and fence. Nevertheless, removing DINOv3 yields a higher Dynamic Degree and overall VBench score.
Metric
VJ
VM
D
D+VJ
VJ+VM
D+VJ+VM
Aggregate scores
Total
83.91
83.97
81.96
83.64
84.64
83.81
Quality
84.62
85.02
82.32
84.26
85.43
84.44
Semantic
81.09
79.79
80.54
81.19
81.48
81.31
Quality dimensions
Subject consistency
96.28
95.89
97.27
96.94
96.58
96.83
Appendix
Table 8 : Detailed encoder ablation on VBench (0–100; higher is better). VJ: V-JEPA 2; VM: VideoMAE; D: DINOv3. Bold indicates the best value within each row, including ties. All variants are trained for 100 steps.
Figure 9 : Alternative video representations. Each row shows sampled frames from a generated video; the left and right panels use InternVideo-Next and Motionformer, respectively. Across the examples, subject detail and scene structure are not reliably maintained. The figure illustrates failure modes under the tested feature losses.
Figure 10 : Adding frame-wise representation losses. Results with the combined SigLIP, Inception, and MAE losses. Columns show sampled frames from each example; red circles mark changes in subject appearance or scene geometry. These examples illustrate temporal inconsistencies that remain despite additional image-level supervision.
Figure 11 : Qualitative ablation of the hybrid reference estimator. Removing either the Monte Carlo or Nyström component introduces visible temporal and structural artifacts (red circles), while the full estimator produces coherent pouring dynamics and stable object geometry.
Figure 12 : Exploratory sparse-feature FD objective. Sampled frames from the tested configuration show a blurred human silhouette with little background detail. This example motivates preserving dense video features but is not a comparison of all FD-based objectives.
Matching space
Total
Quality
Semantic
Text–video joint
83.28
83.73
81.47
Video only
84.05
84.62
81.77
Appendix
Table 9 : Joint text–video versus video-only distribution matching on Wan2.1-T2V-14B. VBench scores are on a 0–100 scale; higher is better.
Figure 13 : FIFO feature reuse. More online samples and gradient compensation improve queued variants, but retain historical features. The top row uses 32 fresh samples without a queue.
Method
TS
VQ
TC
SI
Overall
Monochrome appearance
Self-Forcing
8.822
8.340
8.140
8.106
8.308
Wan2.1-1.3B
8.674
8.390
8.410
8.372
8.382
Wan2.1-14B
9.840
8.638
8.574
8.494
8.992
Elastic Forcing
9.910
8.544
8.526
8.426
8.942
Nailong character
Appendix
Table 10 : Specialized reference-data adaptation. Ratings are on a 0–10 scale; higher is better. TS: task-specific score; VQ: visual quality; TC: temporal consistency; SI: subject integrity. Bold and underlining denote the best and second-best score within each task, respectively.
Figure 14 : Real-video reference data. Generated examples after replacing Wan-generated references with OpenVid videos. Each row shows sampled frames from one generation. The examples span object, human, and landscape content and illustrate learning from a reference source independent of the initialization model’s samples.
Figure 15 : VACE adaptation with visual conditioning. Optical-flow conditions appear on the left and sampled output frames on the right. Training uses 100 MMD-plus-regression updates followed by 100 MMD-only updates, with first-frame and flow conditioning and no separate ODE initialization stage.
Figure 16 : Additional 14B comparisons. Each prompt is followed by Krea Realtime (top) and Elastic Forcing (bottom), with sampled frames arranged from left to right. Red circles highlight inconsistencies in object interaction, rider–bicycle interaction, and bridge geometry in the selected examples.
Figure 17 : Strong agreement between pairwise VLM and human judgments. GPT-5.5 preference margins are compared with human preference differences over 20 matched prompts and 42 raters. The left panel uses raw human scores, while the right panel uses within-rater normalized scores. Positive values favor Ours-14B, and solid lines indicate least-squares fits.
Raw Human Scores
Rater-Normalized Scores
VLM Signal
r
p
r
p
Direct pairwise margin
0.595
0.0057
0.589
0.0063
Appendix
Table 11 : Pairwise VLM–human correlation. Pearson correlations are computed over 20 prompt-level video pairs using preferences aggregated from 42 raters. All p -values are two-sided.
Figure 18 : Additional demonstrations of our 14B model. Each row presents temporally ordered frames from a generated video. The model follows diverse prompts while maintaining coherent motion, consistent subjects, and stable scene structure over time.
Figure 19 : Blind human-evaluation interface. Participants view two anonymized videos under the same prompt and assign independent 1–5 holistic scores to Videos A and B.