On-policy distillation (OPD) has demonstrated two important capabilities in language models: compressing large teachers into smaller students and merging expert models into a single model. Existing diffusion OPD, however, mostly focus on the latter, with teachers and students sharing the same backbone and scale. We investigate large-to-small diffusion opd from large teachers to a small student and find that the standard recipe fails. To find the underlying cause, we propose Fixed-State KL, an effective and fair way to measure the distribution gap between student and teacher during OPD training for diffusion models. We are the first to clarify why large-to-small OPD is challenging for diffusion models: a smaller student struggles to perfectly match the distribution of a larger teacher, while classifier-free guidance can accumulate and amplify the distributional discrepancies between the student's conditional and unconditional branches and those of the teacher. To solve this problem, we propose GFD-OPD, a simple yet effective method that reduces the student-teacher gap while avoiding the error amplification of the CFG composition. Across numerous experiments, GFD outperforms previous baselines in both training efficiency and final performance, achieving state-of-the-art results on all benchmarks.
Figures & tables
Figure 1: GFD-OPD distills SD3.5-Large into SD3.5-Medium. (a) Score versus training compute (GPU hours) on OCR, PickScore, and GenEval: GFD-OPD overtakes all baselines within 10 GPU-hours. (b) GFD-OPD ranks first on all seven evaluation dimensions, covering both task rewards and image quality.
Figure 2
Figure 3: GFD-OPD overview. Left (motivation): the deployed error of a CFG-composed student decomposes as δg=δu+ωd — same-size distillation survives through error cancellation, large-to-small does not, and guidance folding removes the composition so that nothing is left for ω to amplify. Right (method): the process of GFD-OPD.
Figure 4: Shared cross-architecture artifacts under a SD3.5-L teacher. Rows: the SD3.5-M base, three prior OPD recipes ( DiffusionOPD , PDM , splitKL ), GFD and large teacher.
Table 5
Figure 5: Qualitative comparison at 1024×1024 . Columns, left to right: SD3.5-M base; the medium and large teachers; DiffusionOPD M→M ; and GFD L→M .
Method
GenEval
OCR
PickScore
HPSv3
SD3.5-Medium (base)
0.675
0.554
0.843
9.020
Standard OPD
0.884
0.908
0.880
9.175
+ Guidance folding
0.952
0.954
0.923
11.19
+ Extrapolated mean (GFD-OPD)
0.963
0.973
0.930
11.27
Table 5: Ablation: standard OPD + guidance folding + extrapolated-mean target (full GFD-OPD), on the SD3.5-Medium student at 1024×1024 ; HPSv3 probes image quality.
Figure 6: Statistics of the student’s final output as the guidance scale w varies. Each axis is the log ratio of a statistic of the student’s samples to the teacher’s, so the teacher sits at the origin ( ⋆ ). (a) Final latent: low-frequency energy (global tone) vs. high-frequency energy (fine texture). (b) Decoded image: tonal amplitude (RMS contrast × saturation) vs. fine-texture amplitude (3px high-pass MAE). The definition of all metrics are as detailed in Appendix B.3 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Student
∥δu∥
∥d∥
∥r∥
cross term
DiffusionOPD M→M (reference)
0.092
0.032
0.106
−1.56
DiffusionOPD L→M
0.090
0.043
0.177
−0.48
PDM L→M
0.053
0.041
0.178
−0.18
split-KL L→M
0.051
0.043
0.185
−0.18
Appendix
Table 6: Decomposition δg=δu+wd of the guided error at w=4.5 . Cross term =2⟨δu,4.5d⟩/∥r∥2 .
Student
Probe
0–9
10–19
20–29
30–37
0–37 (total)
GFD ( w=1 )
Ptea
0.963
0.984
0.991
0.983
0.980
Preal
0.943
0.971
0.987
0.984
0.968
GFD ( w=2 )
Ptea
1.030
0.991
0.993
0.984
1.000
Preal
1.023
0.978
0.989
0.985
0.997
Appendix
Table 7: Projection g of the student’s deployed field onto the teacher’s guided field.
Student
KL ( Ptea )
KL ( Preal )
GFD L→M ( w=1 )
0.050
0.030
GFD L→M ( w=2 )
0.053
0.032
GFD+uncond L→M ( w=1 )
0.048
0.029
GFD+uncond L→M ( w=2 )
0.060
0.040
Appendix
Table 8: KL divergence to the teacher for GFD and GFD + uncond, in the format of Tab. 2 .
Figure 7: Statistics of the final output of GFD and GFD + uncond as the guidance scale w varies, in the coordinates of Fig. 6 (log ratio to the teacher, teacher at the origin).
Figure 8: Qualitative comparison on FLUX.2 at 1024×1024 , complementing Table 4 . Columns, left to right: the FLUX.2-4B base model; the FLUX.2-9B teacher and the GFD-OPD student distilled from it into 4B; the FLUX.2-32B teacher and the GFD-OPD student distilled from it into 4B. The top four rows are GenEval prompts; the bottom four are LongText prompts probing long-form visual text rendering.
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negative-branch errors can compensate in the guided prediction. Through two contrasting cases, we find that naive matching remains effective under shared negative conditioning, where both branch errors decrease jointly. When the model's native CFG schema retains privileged information in the teacher's negative branch that is unavailable to the student, however, this joint reduction breaks down and the composed objective induces antagonistic branch-error dynamics, reducing the positive-branch error while increasing the negative-branch error. We term this failure mode Negative Branch Asymmetry (NBA). To address NBA, we introduce Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction. We apply PDM to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.
Reinforcement learning has emerged as a powerful tool for improving diffusion-based text-to-image models, but existing methods are largely limited to single-task optimization. Extending RL to multiple tasks is challenging: joint optimization suffers from cross-task interference and imbalance, while cascade RL is cumbersome and prone to catastrophic forgetting. We propose DiffusionOPD, a new multi-task training paradigm for diffusion models based on Online Policy Distillation (OPD). DiffusionOPD first trains task-specific teachers independently, then distills their capabilities into a unified student along the student own rollout trajectories. This decouples single-task exploration from multi-task integration and avoids the optimization burden of solving all tasks jointly from scratch. Theoretically, we lift the OPD framework from discrete tokens to continuous-state Markov processes, deriving a closed-form per-step KL objective that unifies both stochastic SDE and deterministic ODE refinement via mean-matching. We formally and empirically demonstrate that this analytic gradient provides lower variance and better generality compared to conventional PPO-style policy gradients. Extensive experiments show that DiffusionOPD consistently surpasses both multi-reward RL and cascade RL baselines in training efficiency and final performance, while achieving state-of-the-art results on all evaluated benchmarks.
On-policy distillation (OPD) has emerged as an effective post-training paradigm for language models, with recent efforts extending it to block diffusion language models (BDLMs). However, existing studies focus almost exclusively on small block sizes, leaving distillation into student models with larger blocks underexplored. In this work, we investigate this regime and reveal two critical optimization biases that induce severe training instability. First, mismatched block boundaries between teacher and student cause \textbf{\textit{context misalignment}}, providing distorted supervisory signals that misguide student decoding. Second, even under aligned contexts, an \textbf{\textit{intrinsic optimization bias}} in OPD, where the student tends to rapidly absorb high-support signals while lagging on low-support updates, drives a premature confidence surge that traps weaker students in catastrophic overconfidence collapse. To resolve these, we propose \mbox{\textbf{Un-OPD}}, an unbiased on-policy distillation framework with two novelties for stabilizing BDLM training. First, Un-OPD introduces a boundary-aware step filtering strategy that eliminates context-misaligned decoding steps. Second, Un-OPD proposes moderating optimization intensity at high-support positions via a support-rebalanced confidence calibration, thereby bypassing overconfidence collapse. Beyond stability, we also introduce a rollout reuse mechanism to reduce rollout generation overhead. Extensive experiments on math reasoning and code generation benchmarks show that Un-OPD consistently stabilizes training and delivers superior performance while reducing wall-clock training time by approximately half.
Zaiquan Yang, Fei Wei, Yong Wang +6
City University of Hong Kong · Alibaba Group · Beijing Institute of Technology +2