cs.LGDec 28, 2025

ReDiF: Resource-Efficient Few-Step Diffusion Distillation via Reinforcement Learning

Authors: Amirhossein Tighkhorshid, Zahra Dehghanian, Hamid R. Rabiee

Organizations: Computer Engineering Sharif University of Technology

Abstract

Step distillation accelerates diffusion sampling by training a few-step student to imitate a many-step teacher, but distillation itself remains expensive. Typically, this requires thousands of GPU-hours and a large pre-generated trajectory dataset. We introduce ReDiF, which casts step distillation as terminal-reward policy optimization rather than step-wise regression. The student is optimized against a reward computed on the terminal sample, measuring alignment with the teacher's output, instead of matching the teacher's intermediate trajectory under a reconstruction or consistency loss. Because the reward need not be differentiable or trajectory-aligned, ReDiF admits non-differentiable objectives, multi-objective combinations, and preferences the teacher does not express, while exploration lets the student find sampling paths matched to its own step schedule. ReDiF converges in 400 policy updates with 3200 rollouts on a single A100 GPU with 1,000 noise-class pairs and no paired dataset: about 4 GPU-hours, against roughly 336 A100-hours reported for DMD2 on the same EDM teacher. At 8 steps on ImageNet-64, it achieves an FID 3.64 points better than the strongest retrained distillation baseline in the low-training regime under the same single-GPU budget. The formulation is also orthogonal to existing distillation objectives: added to the DMD2 loss, it further improves DMD2's FID by 4.4.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning

    Apr 21, 2026Linwei Dong, Ruoyu Guo, Ge Bai +3Few-Step DistillationDistribution Matching Distillation

  2. Bridging Diffusion Pruning and Step Distillation with Teacher-Aligned Repair

    Jul 7, 2026Jincheng Ying, Li Wenlin, Minghui Xu +1Diffusion ModelsTeacher

  3. AdvDMD: Adversarial Reward Meets DMD For High-Quality Few-Step Generation

    Apr 29, 2026Xu Wang, Zexian Li, Litong Gong +2Few-Step DistillationFew-Step Generation