Unlocking Few-Step Diffusion for Faithful Previews
Authors: Jing Jia, Sifan Liu, Guanyang Wang
Organizations: Department of Computer Science, Rutgers University · Department of Statistical Science, Duke University · Department of Statistics, Rutgers University
Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can closely reproduce their corresponding full-step outputs. Building on this finding, we learn corrections to the initial noise and denoising updates using endpoint supervision, improving correspondence with full-step outputs generated from the same noise and prompt. The resulting previews allow users to screen candidates cheaply and reserve full-step generation for promising ones. Input correction also transfers across sampling budgets without retraining. Experiments show substantial improvements in reference fidelity, including 53-78% lower reconstruction MSE than retrained LD3 on unconditional benchmarks, alongside improved ranking preservation and candidate selection on SD1.5, SDXL, and FLUX.1-dev.
Figures & tables
Figure 1: Traditional vs. preview-then-refinement workflow. Traditional generation requires full denoising steps (e.g. 60 steps) before selection, making rejected candidates expensive. Our enhancement module enables a faithful 3-step preview for cheap selection, with full generation performed only for selected candidates.
Figure 2: Input correction (A) and trajectory correction (B) for fast and faithful diffusion previews.
Figure 3: Noise optimization reveals the reconstruction capacity of frozen few-step samplers. Columns show raw outputs, full-step targets, and noise-optimized outputs.
Model / data
N
few/ref
Raw
Noise-Optimized
Gain (dB)
DDPM / CIFAR-10
100
3 / 1000
17.00±1.91
37.41±1.89
+20.41
DDPM / LSUN-Church
100
3 / 20
14.96±1.49
38.01±3.15
+23.05
DDPM / CelebA-HQ
100
3 / 20
14.62±1.65
39.14±3.89
+24.52
DDPM / LSUN-Bedroom
100
3 / 20
15.52±1.77
39.98±3.09
+24.46
FLUX.1-dev
4
4 / 28
14.54±1.95
37.49±7.69
+22.95
Table 1: Frozen-model oracle noise optimization. PSNR (dB) is the mean ± sample SD over N images; few/ref denotes student/teacher steps.
Figure 4: Representative examples of input correction on CIFAR-10, LSUN-Church, and CelebA-HQ. Columns show uncorrected 3-step DDIM, 4-step DPM-Solver++, 4-step LD3, and our method.
Figure 5: PSNR and HPS on LSUN-Church and CelebA-HQ as the number of DDIM steps increases. Dashed lines are raw DDIMs.
Table 2: Within each dataset/CFG group, metrics are reported in the order MSE ↓ , PSNR ↑ , LPIPS ↓ , and SSIM ↑ . Noise-correction results are averaged over 100 paired test noises per dataset. For input correction, 3+1 denotes three denoiser evaluations and one lighter corrector evaluation. Trajectory correction is evaluated on SDXL and SD1.5. Bold denotes the best reported result within each group.
SDXL Selected ↑ / Regret ↓ / Spearman ↑
Metric
Method
CFG = 1
CFG = 3
CFG = 5
CFG = 7
PickScore
3-step DDIM
19.642
1.377
0.016
21.724
0.915
0.009
22.162
0.855
-0.003
22.312
0.814
-0.011
3-step DPM++
20.385
0.634
0.403
22.152
0.486
0.351
22.435
0.582
0.268
22.423
0.703
0.162
3-step LD3
20.516
0.505
0.476
22.108
0.529
0.368
22.515
0.506
0.280
22.597
0.531
0.223
Ours
20.621
0.397
0.593
22.215
0.424
0.463
22.658
0.359
0.466
22.741
0.386
0.347
ImageReward
3-step DDIM
-1.200
1.242
-0.030
0.274
0.701
-0.022
0.550
0.598
0.003
0.646
0.529
0.021
Table 3: Candidate-ranking preservation on SDXL and SD1.5 across CFG scales. We compare DDIM, DPM++, LD3, and our method, in addition ConSolver for SD1.5. For PickScore, ImageReward, and agent, we report the selected teacher score ↑ , regret ↓ , and Spearman rank correlation ↑ .
Method
ImageReward ↑
PickScore ↑
HPS v2.1 ↑
CLIP ↑
AES ↑
NFE ↓
Random
0.9483
22.9300
29.6458
25.7252
5.6658
28
Diffusion Probe
0.9731
22.9311
29.7302
25.7060
5.6749
70
Probe-Select
0.9442
22.9432
29.6611
25.7982
5.6709
70
ConSolver
1.0085
22.9818
29.8980
25.7976
5.6857
60
TDD-FLUX
1.0331
23.0183
29.9616
25.9744
5.7166
60
Ours
1.1417
23.1869
30.3263
26.4693
5.7625
60
Table 4: FLUX.1-dev selection. HPS v2.1 and CLIP-L/14 cosine: ×100 . Bold: best non-oracle mean. NFE counts screening and final-generation denoiser calls, excluding decoding/scoring.
Table 5: Transfer across steps on 200 samples per dataset. HPS: HPSv2.1 ×100 ; IR: ImageReward. Best values per dataset, K , and metric are bold.
Method
Calls
Params (M)
MSE ↓
PSNR ↑
LPIPS ↓
SSIM ↑
CIFAR-10
Raw DDIM3
3
0
21.18
17.29
0.149
0.578
Input correction
3+1
1.266
2.86
26.60
0.025
0.901
Trajectory correction
3
3.788
1.66
29.79
0.011
0.944
SD1.5, CFG 1
Raw DDIM3
3
0
20.05
17.33
0.801
0.402
Appendix
Table 6: Controlled comparison under matched data, endpoint supervision, and training-update budgets. MSE is multiplied by 103 ; PSNR is in dB. Params counts trainable parameters, and Calls counts U-Net evaluations. The input arm uses one additional full U-Net corrector. References use 1,000 steps for CIFAR-10 and 100 steps for SD1.5. Results use 1,000 CIFAR-10 test inputs or 64 SD1.5 prompts with four noises each.
(a) Preference scores
Method
NFEs
PickScore ↑
HPS ↑
IR ↑
Raw DDIM3
3
19.03±0.45
16.53±2.22
−0.98±0.69
Bedroom-backbone corrector + DDIM3
3+1
20.37±0.54
27.34±1.77
0.54±0.36
CelebA corrector + DDIM3
3+1
19.90±0.62
24.04±2.13
0.10±0.50
DDIM20 teacher
20
19.89±0.61
24.88±2.13
0.10±0.51
(b) Agreement with the DDIM20 teacher
Appendix
Table 7: Cross-backbone transfer from Bedroom to CelebA-HQ. Results are mean ± SD over 200 samples. NFEs denote denoiser plus corrector forward passes. HPS denotes HPSv2.1 ×100 ; IR denotes ImageReward. Preference scores use the fixed text “a portrait photo of a person” for all unconditional outputs.
Figure 7: Comparison of FLUX.1-schnell and our method as previews of FLUX.1-dev. Columns show raw four-step Euler, Schnell, our corrected four-step sampler, and the 28-step reference. Although Schnell achieves higher ImageReward on the previews themselves, our method provides greater fidelity to the reference and better candidate-selection performance.
Figure 8: Noise norm normalization ablation. Columns show raw three-step DDIM, correction without normalization, complete correction, and the 20-step reference.
Figure 9: Correction direction ablation. Columns show raw three-step DDIM, reversed correction, another sample's correction direction rescaled to the original residual norm, complete correction, and the 20-step reference.
Figure 10: Attenuation when transferring a three-step corrector to eight-step DDIM. Columns show raw eight-step DDIM, correction with α=1 , correction with α=(3/8)2.5 , and the 20-step reference.
Figure 11: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, and Ours generations using SDXL at CFG 1.
Figure 12: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, and Ours generations using SDXL at CFG 3.
Figure 13: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, and Ours generations using SDXL at CFG 5.
Figure 14: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, ConSolver, and Ours generations using Stable Diffusion 1.5 at CFG 1.
Figure 15: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, ConSolver, and Ours generations using Stable Diffusion 1.5 at CFG 3.
Figure 16: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, ConSolver, and Ours generations using Stable Diffusion 1.5 at CFG 5.
Figure 17: Comparison of 8 seed for the prompt “A large glass window in a living room.”
Figure 18: Comparison of 8 seed for the prompt “Comfortable, modern living room overlooking a wooded area”.
Figure 19: Comparison of 8 seed selection for the prompt “A living room with windows looking out onto a forest.”
Figure 20: Additional oracle noise examples.
Figure 21: Selected examples of input-correction transfer across sampling steps. Top: examples remaining close to the teacher across step counts. Bottom: examples approaching the teacher as the step count increases. All outputs reuse the same frozen three-step corrector with a fixed exponent of 2.5 and noise norm normalization; each row shares the same original noise.
Figure 22: Additional cross-backbone correction examples on CelebA-HQ. Each triplet shows raw three-step DDIM, the paired 20-step teacher, and three-step DDIM with the transferred input corrector. The corrector was trained through a Bedroom backbone using CelebA teacher targets.
Jul 30, 2026·Hanshuai Cui, Zhiqing Tang, Zhi Yao +3DrafterCorrection
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China