cs.LGSep 30, 2026

GFD-OPD: Guidance-Folded On-Policy Distillation of Diffusion Models Across Scales

Authors: Zhenxing Zhang, Jiayan Teng, Wenxu Wu, Zhuoyi Yang, Jiazheng Xu, Wendi Zheng, Jie Tang, Dan Guo, +1 more

Organizations: Hefei University of Technology · Zhipu AI · Tsinghua University

Abstract

On-policy distillation (OPD) has demonstrated two important capabilities in language models: compressing large teachers into smaller students and merging expert models into a single model. Existing diffusion OPD, however, mostly focus on the latter, with teachers and students sharing the same backbone and scale. We investigate large-to-small diffusion opd from large teachers to a small student and find that the standard recipe fails. To find the underlying cause, we propose Fixed-State KL, an effective and fair way to measure the distribution gap between student and teacher during OPD training for diffusion models. We are the first to clarify why large-to-small OPD is challenging for diffusion models: a smaller student struggles to perfectly match the distribution of a larger teacher, while classifier-free guidance can accumulate and amplify the distributional discrepancies between the student's conditional and unconditional branches and those of the teacher. To solve this problem, we propose GFD-OPD, a simple yet effective method that reduces the student-teacher gap while avoiding the error amplification of the CFG composition. Across numerous experiments, GFD outperforms previous baselines in both training efficiency and final performance, achieving state-of-the-art results on all benchmarks.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation

    Jul 27, 2026Bingnan Li, Haozhe Wang, Haozhong Xiong +5Classifier-Free GuidancePositive

  2. Towards Unbiased On-Policy Distillation for Block Diffusion Language Models

    Oct 4, 2026Zaiquan Yang, Fei Wei, Yong Wang +6