On-policy distillation (OPD) specializes pretrained video diffusion models through teacher supervision along the student's own generation trajectory. Although large video models are natural teachers, developing specialized video experts can require costly video data and training, while querying them incurs substantially higher latency than querying image experts. More readily available and cheaper to query, image experts offer a cost-effective alternative, particularly for largely temporal-agnostic capabilities such as aesthetics and OCR that admit frame-level supervision. However, heterogeneous image and video latent spaces prevent direct supervision of intermediate student states, while image experts lack cross-frame motion supervision, making temporal consistency vulnerable to frame-level improvements. In this paper, we propose MILD, a Motion-Preserving Image-to-Video Latent Distillation framework that transfers specialized image expertise while preserving pretrained video dynamics. MILD uses a learnable linear connector that aligns student latent states and predicted updates with those of image experts, enabling supervision transfer across heterogeneous latent spaces. We further constrain image-guided corrections around the pretrained student's predictions to preserve video dynamics and incorporate an optical-flow-based motion reward to improve motion quality and temporal consistency. Across specialized image experts and multiple video-student backbones, our method consistently outperforms video-teacher OPD baselines, with further studies demonstrating effective transfer across connector designs and heterogeneous architectures. These results establish image-to-video distillation as an effective route to improving video generation by drawing on the diverse and evolving capabilities of the image-generation ecosystem.
Figures & tables
Figure 1: Computational costs of I2V and V2V OPD. Both train Wan2.1-1.3B using SD3.5-Medium and Wan2.1-14B teachers, respectively. Left: estimated student training-state and teacher memory, excluding activations and auxiliary modules. Right: measured teacher-query latency.
Figure 2: Qualitative comparisons of Base, V2V OPD, and I2V OPD. Base denotes the pretrained Wan2.1-1.3B model. Top: aesthetic quality on DrawBench, text rendering on the Text Rendering benchmark, and compositional generation. Bottom: complex scene generation, count and action accuracy, and camera motion control on VBench-2.0 prompts.
Figure 3: Overview of MILD. Learnable connectors transfer image-expert supervision to on-policy video states. A frozen video anchor bounds spatial corrections, while optical-flow feedback provides motion refinement. Predictions are fused into a detached distillation target; connectors use a separate alignment objective.
Video student
V2V teacher
Image experts
Wan2.1-1.3B
Wan2.1-14B
SD3.5-Medium
Wan2.2-5B
Wan2.2-A14B
Wan2.2-5B
Wan2.2-A14B
LTX-Video-2B
LTX-Video-13B
Table 1: Experimental setup. Video-teacher pairings and shared image-expert backbone. Each V2V configuration uses one video teacher.
VBench-2.0 ↑
EvalCrafter ↑
Method
Creativity
Commonsense
Controllability
Human Fidelity
Physics
AVG.
Visual Quality
Text–Video Alignment
Motion Quality
Temporal Consistency
Final Sum Score
Large video teachers (reference)
Wan2.1-14B
48.91
59.95
27.44
90.94
44.09
54.27
66.25
56.78
53.97
63.93
240.93
Wan2.2-A14B
52.98
67.73
35.36
82.59
51.60
58.05
66.47
60.22
54.58
63.10
244.37
LTX-Video-13B
55.02
40.84
13.90
87.60
43.14
48.10
64.08
51.99
56.00
65.04
237.11
Wan2.1-1.3B
Table 2: Evaluation results on VBench-2.0 and EvalCrafter. MILD outperforms both V2V OPD baselines on the aggregate scores of both benchmarks across all three student backbones. Compared with the larger-teacher V2V OPD baseline, MILD reduces the sum of per-GPU peak allocated memory by 39.32%, 64.07%, and 19.53% on Wan2.1-1.3B, Wan2.2-5B, and LTX-Video-2B, respectively (Figure 5 ). Bold and underlined values indicate the best and second-best results within each student group.
Figure 4: Spatial capability comparisons across Wan2.1 and Wan2.2 settings. Each setting compares the student, a larger video teacher, a video teacher post-trained on aesthetic video data, and specialized image teachers. Image experts achieve the highest GenEval, OCR, and Aesthetic scores in both settings.
Method
Creativity
Commonsense
Controllability
Human Fidelity
Physics
AVG.
Wan2.1-1.3B
45.86
59.68
22.85
85.16
42.07
51.12
V2V OPD (Wan2.1-14B)
47.34
63.71
24.67
85.17
45.32
53.24
V2V OPD (Wan2.2-A14B)
48.74
65.72
24.59
82.37
42.91
52.87
V2V OPD (Wan2.2-5B)
48.29
58.50
23.70
82.51
43.14
51.23
Table 3: Cross-architecture video-to-video OPD on VBench-2.0. All OPD variants use Wan2.1-1.3B as the student, with the teacher specified in parentheses. Distillation from Wan2.2 teachers improves the student’s overall performance, showing that the proposed connector transfers supervision between heterogeneous video diffusion models.
Figure 5: GPU memory comparison across student backbones. MILD reduces the sum of per-GPU peak allocated memory relative to V2V OPD across all three backbones. The metric sums each participating GPU’s peak allocated memory.
Figure 6: Qualitative comparisons using Wan2.2-5B as the student. Frames 0, 44, and 80 illustrate MILD’s ability to accurately follow prompts involving color changes and directional motion. In contrast, baselines fail to respond correctly to the moon’s color change and produces artifacts (left), while the cat’s motion direction is incorrect (right). More visualizations are in Appendix C.3 .
DrawBench
FlowGRPO splits
Method
Aesthetic
PickScore
HPSv2
CLIPScore
ImageReward
GenEval
OCR
Base
4.7855
0.7841
0.2293
0.2372
−0.4390
0.3252
0.2522
V2V OPD (Aes. SFT teacher)
4.9169
0.7931
0.2321
0.2414
−0.3098
0.4023
0.2617
V2V OPD (Wan2.2-A14B)
4.8217
0.7819
0.2253
0.2338
−0.4610
0.4067
0.2556
MILD
5.0305
0.8111
0.2470
0.2438
0.0320
0.4454
0.2697
Table 4: Frame-level spatial capability evaluation using Wan2.2-5B. We evaluate the middle frame of each generated video, using DrawBench prompts for the five model-based metrics and FlowGRPO evaluation splits for GenEval and OCR, consistent with DiffusionOPD ( Li et al., 2026 ) . MILD achieves the highest score on every metric among the compared student variants.
Variant
Creativity
Commonsense
Controllability
Human Fidelity
Physics
AVG.
Pretrained model (reference)
Wan2.2-5B (Base)
48.20
58.50
20.02
81.89
49.31
51.58
Supervision and optimization strategy
w/o image experts
45.97
60.79
20.48
82.11
50.95
52.06
w/o motion refinement
48.64
61.75
19.73
83.45
53.68
53.45
V2V OPD (Wan2.2-A14B) + motion refinement
46.86
61.68
20.40
83.61
52.74
53.06
Table 5: Ablation study on Wan2.2-5B using VBench-2.0. The full model uses auxiliary motion refinement, joint optimization, the default connector, and prompt-dependent image-score weights that enable OCR guidance only for text-related prompts. Single-expert variants use only the indicated image expert.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Wan2.1-1.3B
Wan2.2-5B
LTX-Video-2B
Student denoising steps
50
50
40
OPD trajectory-index range
2–47
2–47
2–37
Student learning rate
5×10−7
3×10−7
2×10−6
Connector learning rate
5×10−6
1×10−5
2×10−6
Connector hidden channels
16
48
128
Image-supervision temporal stride
4
4
8
Appendix
Table 6: Backbone-specific implementation settings. Trajectory-index ranges specify candidate states for OPD supervision; temporal stride is measured in RGB-frame indices.
Source
Train
Validation
Real videos
3,412
379
Generated videos
3,412
379
Temporal hard negatives
3,412
379
Total
10,236
1,137
Appendix
Table 7: Motion-discriminator training data. Generated negatives use the corresponding video backbone; counts precede balanced sampling.
Student backbone
AUROC
Pairwise accuracy
Reward–magnitude correlation
Wan2.1-1.3B
0.9552
0.9472
−0.0817
Wan2.2-5B
0.9288
0.9736
−0.1078
LTX-Video-2B
0.8925
0.9538
−0.0521
Appendix
Table 8: Validation of the motion discriminators. Higher AUROC and pairwise accuracy are better. Reward–magnitude correlation is reported as a signed value.
VBench-2.0
EvalCrafter
Method
Creativity
Commonsense
Controllability
Human Fidelity
Physics
AVG.
Visual Quality
Text–Video Alignment
Motion Quality
Temporal Consistency
Final Sum Score
Wan2.1-1.3B
Aesthetic SFT
49.84
61.13
24.94
86.45
44.94
53.46
66.38
58.90
54.18
63.34
242.80
MILD
47.89
64.56
25.17
85.47
46.52
53.92
66.84
57.89
54.58
63.79
243.10
Wan2.2-5B
Aesthetic SFT
50.65
63.13
22.94
83.44
46.76
53.38
63.60
57.89
54.45
62.88
238.82
Appendix
Table 9: Comparison with direct aesthetic video fine-tuning. Aesthetic SFT denotes video models fine-tuned on aesthetic video data and evaluated directly, without subsequent OPD. Bold indicates the better result within each backbone. Higher is better for all metrics.
Figure 7: Qualitative comparison with direct aesthetic video fine-tuning on VBench-2.0 using Wan2.2-5B. Frames 0, 44, and 80 illustrate appearance changes and directional motion under matched prompts.
Figure 8: Abandoned-city and underground-cave comparisons. Left: MILD depicts prominent illuminated signs, wall markings, and wet-pavement reflections. Right: MILD renders distinct foreground stalactites, textured rock surfaces, and visible bat silhouettes. Each example shows frames 0, 44, and 80.
Figure 9: Text-rendering and twilight-city comparisons. Left: MILD renders “START” more clearly in the frames where the complete word is visible. Right: MILD depicts distinct architecture, illuminated bridge contours, and water reflections. Each example shows frames 0, 44, and 80.
Figure 10: Whale-and-seagull composition. MILD depicts a more prominent seagull alongside the whale, with both subjects clearly identifiable across frames 0, 44, and 80.
Figure 11: Qualitative comparison of aesthetic image and video teachers. Each pair shows the Wan2.2-5B video teacher post-trained on aesthetic video data (left) and the aesthetic image expert (right) under the same prompt. Rows show DrawBench, GenEval, and OCR examples, illustrating differences in visual detail, compositional accuracy, and text rendering available for distillation.
Figure 12: Spatial capability comparisons using Wan2.2-5B. Columns show the pretrained student, OPD with an aesthetic-SFT video teacher, OPD with a larger video teacher, and MILD. Matched-prompt examples illustrate color-attribute alignment, text rendering, and object composition in sampled video frames.
Figure 13: Spatial capability comparison with direct aesthetic video fine-tuning using Wan2.2-5B. Matched-prompt examples illustrate color-attribute alignment, text rendering, and object composition in sampled video frames.