ReDiF: Resource-Efficient Few-Step Diffusion Distillation via Reinforcement Learning
Organizations: Computer Engineering Sharif University of Technology
Abstract
Step distillation accelerates diffusion sampling by training a few-step student to imitate a many-step teacher, but distillation itself remains expensive. Typically, this requires thousands of GPU-hours and a large pre-generated trajectory dataset. We introduce ReDiF, which casts step distillation as terminal-reward policy optimization rather than step-wise regression. The student is optimized against a reward computed on the terminal sample, measuring alignment with the teacher's output, instead of matching the teacher's intermediate trajectory under a reconstruction or consistency loss. Because the reward need not be differentiable or trajectory-aligned, ReDiF admits non-differentiable objectives, multi-objective combinations, and preferences the teacher does not express, while exploration lets the student find sampling paths matched to its own step schedule. ReDiF converges in 400 policy updates with 3200 rollouts on a single A100 GPU with 1,000 noise-class pairs and no paired dataset: about 4 GPU-hours, against roughly 336 A100-hours reported for DMD2 on the same EDM teacher. At 8 steps on ImageNet-64, it achieves an FID 3.64 points better than the strongest retrained distillation baseline in the low-training regime under the same single-GPU budget. The formulation is also orthogonal to existing distillation objectives: added to the DMD2 loss, it further improves DMD2's FID by 4.4.
Figures & tables
| Method | Steps | Batch | GPU-h | PFLOPs | P | R | D | C | FID | Pub. FID |
| (a) EDM teacher, ImageNet-64 ( ) | ||||||||||
| Teacher (256 steps, reference) | 256 | – | – | – | 0.8310 | 0.7103 | 0.8930 | 0.7831 | 13.74 | 1.36 |
| Teacher (8 steps) | 8 | – | – | – | 0.5308 | 0.4780 | 0.3910 | 0.5326 | 35.04 | – |
| Progressive Distill. ( Salimans and Ho, 2022 ) | 32 | 15.39 | ||||||||
| Consistency Distill. ( Song et al., 2023 ) | 32 | 6.20 | ||||||||
| DMD ( Yin et al., 2024b ) | 128 | 2.62 | ||||||||
| Method | Mode | Aux. net | PFLOPs | P | R | D | C | FID |
|---|---|---|---|---|---|---|---|---|
| Progressive Distill. ( Salimans and Ho, 2022 ) | – | interm. teachers | ||||||
| + RL (ours) | structural | interm. teachers | ||||||
| Consistency Distill. ( Song et al., 2023 ) | – | EMA target | ||||||
| + RL (ours) | structural | EMA target | ||||||
| DMD2 ( Yin et al., 2024a ) | – | score net + disc. | ||||||
| + RL (ours) | loss | score net + disc. |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Step Distill. | One-Step | Data-Free | RL | Few Pr. | Non-Diff. | Low-Comp. |
|---|---|---|---|---|---|---|---|
| Trajectory matching | |||||||
| Progressive Distill. ( Salimans and Ho, 2022 ) | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Guided Distill. ( Meng et al., 2023 ) | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Consistency ( Song et al., 2023 ) | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| LCM / LCM-LoRA ( Luo et al., 2023a ; Luo et al., 2023b ) | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| CTM ( Kim et al., 2024 ) | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Hyperparameter | EDM / ImageNet-64 | SD v1.5 / COCO |
|---|---|---|
| Optimizer | AdamW | AdamW |
| Trainable parameters | LoRA, rank / alpha 4 / 4 | LoRA, rank / alpha 4 / 4 |
| LoRA target modules | UNet attention ( ) | UNet attention ( ) |
| Learning rate | ||
| PPO clip range | ||
| Inner epochs per rollout | 4 | 4 |
| Reward | P | R | D | C | CLIP-T | FID |
|---|---|---|---|---|---|---|
| Teacher (256 steps, reference) | 0.8310 | 0.7103 | 0.8930 | 0.7831 | – | 13.74 |
| Teacher (8 steps) | 0.5308 | 0.4780 | 0.3910 | 0.5326 | – | 35.04 |
| Single-encoder semantic rewards | ||||||
| CLIP (R 1 ) | ||||||
| DINOv3 | ||||||
| PE | ||||||
| Reward | P | R | D | C | FID | CLIPScore |
|---|---|---|---|---|---|---|
| Teacher (50 steps, reference) | 0.6103 | 0.2029 | 0.8878 | 0.6832 | 24.59 | 0.3137 |
| Teacher (5 steps) | 0.4420 | 0.1762 | 0.6291 | 0.5813 | 31.45 | 0.3041 |
| Single-encoder semantic rewards | ||||||
| CLIP (R 1 ) | ||||||
| DINOv3 | ||||||
| PE | ||||||
| Student | LR | P | R | D | C | FID | |
|---|---|---|---|---|---|---|---|
| LoRA ( ) | |||||||
| LoRA ( ) | |||||||
| Full | |||||||
| Full |
| Method | Family | Steps | Aux. net | Trainable | Batch | Devices / training cost (reported) | FID |
|---|---|---|---|---|---|---|---|
| (a) ImageNet-64 | |||||||
| EDM teacher ( Karras et al., 2022 ) | – | 256 | – | – | – | – | 1.36 |
| Progressive Distill. ( Salimans and Ho, 2022 ) | trajectory | interm. teachers | full | 2048 | 64 GPUs; 5,533 A100-h † | 15.39 ‡ | |
| Consistency Distill. ( Song et al., 2023 ) | trajectory | target copy | full | 2048 | 64 GPUs; 7,867 A100-h † | 6.20 ‡ | |
| CTM ( Kim et al., 2024 ) | trajectory | target copy + disc. | full | 2048 | 64 GPUs; 902 A100-h † | 1.92 | |
| DMD ( Yin et al., 2024b ) | distribution | score net | full | n/r | n/r | 2.62 | |
| Branch | Method | LR | Aux. LR | Target type |
| ImageNet-64 | Progressive Distillation | – | fixed teacher target | |
| Consistency Distillation | – | fixed teacher target | ||
| DMD | moving target (aux. network) | |||
| DMD2 | moving target (aux. network) | |||
| SiD | moving target (aux. network) | |||
| SiDA | moving target (aux. network) |