On-policy distillation (OPD) has demonstrated two important capabilities in language models: compressing large teachers into smaller students and merging expert models into a single model. Existing diffusion OPD, however, mostly focus on the latter, with teachers and students sharing the same backbone and scale. We investigate large-to-small diffusion opd from large teachers to a small student and find that the standard recipe fails. To find the underlying cause, we propose Fixed-State KL, an effective and fair way to measure the distribution gap between student and teacher during OPD training for diffusion models. We are the first to clarify why large-to-small OPD is challenging for diffusion models: a smaller student struggles to perfectly match the distribution of a larger teacher, while classifier-free guidance can accumulate and amplify the distributional discrepancies between the student's conditional and unconditional branches and those of the teacher. To solve this problem, we propose GFD-OPD, a simple yet effective method that reduces the student-teacher gap while avoiding the error amplification of the CFG composition. Across numerous experiments, GFD outperforms previous baselines in both training efficiency and final performance, achieving state-of-the-art results on all benchmarks.
Figures & tables
Figure 1: GFD-OPD distills SD3.5-Large into SD3.5-Medium. (a) Score versus training compute (GPU hours) on OCR, PickScore, and GenEval: GFD-OPD overtakes all baselines within 10 GPU-hours. (b) GFD-OPD ranks first on all seven evaluation dimensions, covering both task rewards and image quality.
Figure 2
Figure 3: GFD-OPD overview. Left (motivation): the deployed error of a CFG-composed student decomposes as δg=δu+ωd — same-size distillation survives through error cancellation, large-to-small does not, and guidance folding removes the composition so that nothing is left for ω to amplify. Right (method): the process of GFD-OPD.
Figure 4: Shared cross-architecture artifacts under a SD3.5-L teacher. Rows: the SD3.5-M base, three prior OPD recipes ( DiffusionOPD , PDM , splitKL ), GFD and large teacher.
Table 5
Figure 5: Qualitative comparison at 1024×1024 . Columns, left to right: SD3.5-M base; the medium and large teachers; DiffusionOPD M→M ; and GFD L→M .
Method
GenEval
OCR
PickScore
HPSv3
SD3.5-Medium (base)
0.675
0.554
0.843
9.020
Standard OPD
0.884
0.908
0.880
9.175
+ Guidance folding
0.952
0.954
0.923
11.19
+ Extrapolated mean (GFD-OPD)
0.963
0.973
0.930
11.27
Table 5: Ablation: standard OPD + guidance folding + extrapolated-mean target (full GFD-OPD), on the SD3.5-Medium student at 1024×1024 ; HPSv3 probes image quality.
Figure 6: Statistics of the student’s final output as the guidance scale w varies. Each axis is the log ratio of a statistic of the student’s samples to the teacher’s, so the teacher sits at the origin ( ⋆ ). (a) Final latent: low-frequency energy (global tone) vs. high-frequency energy (fine texture). (b) Decoded image: tonal amplitude (RMS contrast × saturation) vs. fine-texture amplitude (3px high-pass MAE). The definition of all metrics are as detailed in Appendix B.3 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Student
∥δu∥
∥d∥
∥r∥
cross term
DiffusionOPD M→M (reference)
0.092
0.032
0.106
−1.56
DiffusionOPD L→M
0.090
0.043
0.177
−0.48
PDM L→M
0.053
0.041
0.178
−0.18
split-KL L→M
0.051
0.043
0.185
−0.18
Appendix
Table 6: Decomposition δg=δu+wd of the guided error at w=4.5 . Cross term =2⟨δu,4.5d⟩/∥r∥2 .
Student
Probe
0–9
10–19
20–29
30–37
0–37 (total)
GFD ( w=1 )
Ptea
0.963
0.984
0.991
0.983
0.980
Preal
0.943
0.971
0.987
0.984
0.968
GFD ( w=2 )
Ptea
1.030
0.991
0.993
0.984
1.000
Preal
1.023
0.978
0.989
0.985
0.997
Appendix
Table 7: Projection g of the student’s deployed field onto the teacher’s guided field.
Student
KL ( Ptea )
KL ( Preal )
GFD L→M ( w=1 )
0.050
0.030
GFD L→M ( w=2 )
0.053
0.032
GFD+uncond L→M ( w=1 )
0.048
0.029
GFD+uncond L→M ( w=2 )
0.060
0.040
Appendix
Table 8: KL divergence to the teacher for GFD and GFD + uncond, in the format of Tab. 2 .
Figure 7: Statistics of the final output of GFD and GFD + uncond as the guidance scale w varies, in the coordinates of Fig. 6 (log ratio to the teacher, teacher at the origin).
Figure 8: Qualitative comparison on FLUX.2 at 1024×1024 , complementing Table 4 . Columns, left to right: the FLUX.2-4B base model; the FLUX.2-9B teacher and the GFD-OPD student distilled from it into 4B; the FLUX.2-32B teacher and the GFD-OPD student distilled from it into 4B. The top four rows are GenEval prompts; the bottom four are LongText prompts probing long-form visual text rendering.