cs.CVSep 28, 2026

From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models

Authors: Bingqing Jiang, Li Luo, Zichao Yu, Yujin Han, Zhaolong Su, Difan Zou

Organizations: The University of Hong Kong · Cornell University

Abstract

On-policy distillation (OPD) specializes pretrained video diffusion models through teacher supervision along the student's own generation trajectory. Although large video models are natural teachers, developing specialized video experts can require costly video data and training, while querying them incurs substantially higher latency than querying image experts. More readily available and cheaper to query, image experts offer a cost-effective alternative, particularly for largely temporal-agnostic capabilities such as aesthetics and OCR that admit frame-level supervision. However, heterogeneous image and video latent spaces prevent direct supervision of intermediate student states, while image experts lack cross-frame motion supervision, making temporal consistency vulnerable to frame-level improvements. In this paper, we propose MILD, a Motion-Preserving Image-to-Video Latent Distillation framework that transfers specialized image expertise while preserving pretrained video dynamics. MILD uses a learnable linear connector that aligns student latent states and predicted updates with those of image experts, enabling supervision transfer across heterogeneous latent spaces. We further constrain image-guided corrections around the pretrained student's predictions to preserve video dynamics and incorporate an optical-flow-based motion reward to improve motion quality and temporal consistency. Across specialized image experts and multiple video-student backbones, our method consistently outperforms video-teacher OPD baselines, with further studies demonstrating effective transfer across connector designs and heterogeneous architectures. These results establish image-to-video distillation as an effective route to improving video generation by drawing on the diverse and evolving capabilities of the image-generation ecosystem.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation

    Jun 20, 2026Wenhu Zhang, Kun Cheng, Changyuan Wang +7Video Diffusion ModelsDistribution Matching Distillation

  2. SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation

    May 28, 2026Zhuguanyu Wu, Ruihao Gong, Yang Yong +5Video Diffusion ModelsAutoregressive Video Diffusion Models

  3. Parallel Decoding Distillation for Fast Image and Video Generation

    Jul 28, 2026Neta Shaul, Chao Liu, Arash Vahdat +1Few-Step DistillationDataset Distillation