Unlocking Few-Step Diffusion for Faithful Previews
Authors: Jing Jia, Sifan Liu, Guanyang Wang
Organizations: Department of Computer Science, Rutgers University · Department of Statistical Science, Duke University · Department of Statistics, Rutgers University
Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can closely reproduce their corresponding full-step outputs. Building on this finding, we learn corrections to the initial noise and denoising updates using endpoint supervision, improving correspondence with full-step outputs generated from the same noise and prompt. The resulting previews allow users to screen candidates cheaply and reserve full-step generation for promising ones. Input correction also transfers across sampling budgets without retraining. Experiments show substantial improvements in reference fidelity, including 53-78% lower reconstruction MSE than retrained LD3 on unconditional benchmarks, alongside improved ranking preservation and candidate selection on SD1.5, SDXL, and FLUX.1-dev.
Figures & tables
Figure 1: Traditional vs. preview-then-refinement workflow. Traditional generation requires full denoising steps (e.g. 60 steps) before selection, making rejected candidates expensive. Our enhancement module enables a faithful 3-step preview for cheap selection, with full generation performed only for selected candidates.
Figure 2: Input correction (A) and trajectory correction (B) for fast and faithful diffusion previews.
Figure 3: Noise optimization reveals the reconstruction capacity of frozen few-step samplers. Columns show raw outputs, full-step targets, and noise-optimized outputs.
Model / data
N
few/ref
Raw
Noise-Optimized
Gain (dB)
DDPM / CIFAR-10
100
3 / 1000
17.00±1.91
37.41±1.89
+20.41
DDPM / LSUN-Church
100
3 / 20
14.96±1.49
38.01±3.15
+23.05
DDPM / CelebA-HQ
100
3 / 20
14.62±1.65
39.14±3.89
+24.52
DDPM / LSUN-Bedroom
100
3 / 20
15.52±1.77
39.98±3.09
+24.46
FLUX.1-dev
4
4 / 28
14.54±1.95
37.49±7.69
+22.95
Table 1: Frozen-model oracle noise optimization. PSNR (dB) is the mean ± sample SD over N images; few/ref denotes student/teacher steps.
Figure 4: Representative examples of input correction on CIFAR-10, LSUN-Church, and CelebA-HQ. Columns show uncorrected 3-step DDIM, 4-step DPM-Solver++, 4-step LD3, and our method.
Figure 5: PSNR and HPS on LSUN-Church and CelebA-HQ as the number of DDIM steps increases. Dashed lines are raw DDIMs.
Table 2: Within each dataset/CFG group, metrics are reported in the order MSE ↓ , PSNR ↑ , LPIPS ↓ , and SSIM ↑ . Noise-correction results are averaged over 100 paired test noises per dataset. For input correction, 3+1 denotes three denoiser evaluations and one lighter corrector evaluation. Trajectory correction is evaluated on SDXL and SD1.5. Bold denotes the best reported result within each group.
SDXL Selected ↑ / Regret ↓ / Spearman ↑
Metric
Method
CFG = 1
CFG = 3
CFG = 5
CFG = 7
PickScore
3-step DDIM
19.642
1.377
0.016
21.724
0.915
0.009
22.162
0.855
-0.003
22.312
0.814
-0.011
3-step DPM++
20.385
0.634
0.403
22.152
0.486
0.351
22.435
0.582
0.268
22.423
0.703
0.162
3-step LD3
20.516
0.505
0.476
22.108
0.529
0.368
22.515
0.506
0.280
22.597
0.531
0.223
Ours
20.621
0.397
0.593
22.215
0.424
0.463
22.658
0.359
0.466
22.741
0.386
0.347
ImageReward
3-step DDIM
-1.200
1.242
-0.030
0.274
0.701
-0.022
0.550
0.598
0.003
0.646
0.529
0.021
Table 3: Candidate-ranking preservation on SDXL and SD1.5 across CFG scales. We compare DDIM, DPM++, LD3, and our method, in addition ConSolver for SD1.5. For PickScore, ImageReward, and agent, we report the selected teacher score ↑ , regret ↓ , and Spearman rank correlation ↑ .
Method
ImageReward ↑
PickScore ↑
HPS v2.1 ↑
CLIP ↑
AES ↑
NFE ↓
Random
0.9483
22.9300
29.6458
25.7252
5.6658
28
Diffusion Probe
0.9731
22.9311
29.7302
25.7060
5.6749
70
Probe-Select
0.9442
22.9432
29.6611
25.7982
5.6709
70
ConSolver
1.0085
22.9818
29.8980
25.7976
5.6857
60
TDD-FLUX
1.0331
23.0183
29.9616
25.9744
5.7166
60
Ours
1.1417
23.1869
30.3263
26.4693
5.7625
60
Table 4: FLUX.1-dev selection. HPS v2.1 and CLIP-L/14 cosine: ×100 . Bold: best non-oracle mean. NFE counts screening and final-generation denoiser calls, excluding decoding/scoring.
Table 5: Transfer across steps on 200 samples per dataset. HPS: HPSv2.1 ×100 ; IR: ImageReward. Best values per dataset, K , and metric are bold.
Method
Calls
Params (M)
MSE ↓
PSNR ↑
LPIPS ↓
SSIM ↑
CIFAR-10
Raw DDIM3
3
0
21.18
17.29
0.149
0.578
Input correction
3+1
1.266
2.86
26.60
0.025
0.901
Trajectory correction
3
3.788
1.66
29.79
0.011
0.944
SD1.5, CFG 1
Raw DDIM3
3
0
20.05
17.33
0.801
0.402
Appendix
Table 6: Controlled comparison under matched data, endpoint supervision, and training-update budgets. MSE is multiplied by 103 ; PSNR is in dB. Params counts trainable parameters, and Calls counts U-Net evaluations. The input arm uses one additional full U-Net corrector. References use 1,000 steps for CIFAR-10 and 100 steps for SD1.5. Results use 1,000 CIFAR-10 test inputs or 64 SD1.5 prompts with four noises each.
(a) Preference scores
Method
NFEs
PickScore ↑
HPS ↑
IR ↑
Raw DDIM3
3
19.03±0.45
16.53±2.22
−0.98±0.69
Bedroom-backbone corrector + DDIM3
3+1
20.37±0.54
27.34±1.77
0.54±0.36
CelebA corrector + DDIM3
3+1
19.90±0.62
24.04±2.13
0.10±0.50
DDIM20 teacher
20
19.89±0.61
24.88±2.13
0.10±0.51
(b) Agreement with the DDIM20 teacher
Appendix
Table 7: Cross-backbone transfer from Bedroom to CelebA-HQ. Results are mean ± SD over 200 samples. NFEs denote denoiser plus corrector forward passes. HPS denotes HPSv2.1 ×100 ; IR denotes ImageReward. Preference scores use the fixed text “a portrait photo of a person” for all unconditional outputs.
Figure 7: Comparison of FLUX.1-schnell and our method as previews of FLUX.1-dev. Columns show raw four-step Euler, Schnell, our corrected four-step sampler, and the 28-step reference. Although Schnell achieves higher ImageReward on the previews themselves, our method provides greater fidelity to the reference and better candidate-selection performance.
Figure 8: Noise norm normalization ablation. Columns show raw three-step DDIM, correction without normalization, complete correction, and the 20-step reference.
Figure 9: Correction direction ablation. Columns show raw three-step DDIM, reversed correction, another sample's correction direction rescaled to the original residual norm, complete correction, and the 20-step reference.
Figure 10: Attenuation when transferring a three-step corrector to eight-step DDIM. Columns show raw eight-step DDIM, correction with α=1 , correction with α=(3/8)2.5 , and the 20-step reference.
Figure 11: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, and Ours generations using SDXL at CFG 1.
Figure 12: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, and Ours generations using SDXL at CFG 3.
Figure 13: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, and Ours generations using SDXL at CFG 5.
Figure 14: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, ConSolver, and Ours generations using Stable Diffusion 1.5 at CFG 1.
Figure 15: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, ConSolver, and Ours generations using Stable Diffusion 1.5 at CFG 3.
Figure 16: Visual comparison of 60-step Reference images and 3-step DDIM, DPM++, LD3, ConSolver, and Ours generations using Stable Diffusion 1.5 at CFG 5.
Figure 17: Comparison of 8 seed for the prompt “A large glass window in a living room.”
Figure 18: Comparison of 8 seed for the prompt “Comfortable, modern living room overlooking a wooded area”.
Figure 19: Comparison of 8 seed selection for the prompt “A living room with windows looking out onto a forest.”
Figure 20: Additional oracle noise examples.
Figure 21: Selected examples of input-correction transfer across sampling steps. Top: examples remaining close to the teacher across step counts. Bottom: examples approaching the teacher as the step count increases. All outputs reuse the same frozen three-step corrector with a fixed exponent of 2.5 and noise norm normalization; each row shares the same original noise.
Figure 22: Additional cross-backbone correction examples on CelebA-HQ. Each triplet shows raw three-step DDIM, the paired 20-step teacher, and three-step DDIM with the transferred input corrector. The corrector was trained through a Bedroom backbone using CelebA teacher targets.
Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. At selected denoising checkpoints, PreviewDiff decodes a partial preview, asks a multimodal judge to score and critique it, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations. These branches are then scored and selectively rolled forward, allowing verifier compute to guide generation while the sample is still editable. Across image and video generation benchmarks, PreviewDiff consistently improves over budget-matched Best-of-N selection and strong scalar-search baselines. Ablations show that earlier interventions and increased search width provide the largest gains, while deeper search and additional semantic variants offer complementary improvements. PreviewDiff demonstrates that multimodal feedback is most useful not only as a final verifier, but as an active controller inside the denoising process.
Diffusion language models (DLMs) promise fast parallel generation, yet high-quality samples often require large number of refinement steps, which diminishes their advantage in practice. This has led to massive interest in and rapid development of new methods for effective few-step generation. We show that much of the supposed quality gap at few steps can instead arise from a suboptimally configured sampler. Modest sampler sharpening, without any model retraining, enables a couple years old masked DLM to rival supposedly far improved successors. This differently sampled DLM in fact achieves lower generative perplexity in just 16 steps than what its standard sampler obtains with 1024, while improving both judged quality and semantic diversity. We further show that conventional per-output metrics can fundamentally obscure these gains, since any optimal trade-off between two such metrics can be attained by a generator supported on at most two outputs. We subsequently introduce GroupEval, which separately evaluates quality and across-output semantic diversity, and offers fresh insights including uncovering how 1.5-4.7x perplexity gains of a distilled model yield no corresponding quality gain. Finally, we explain why sharpening helps: parallel unmasking destroys dependencies among simultaneously generated tokens, creating a gap between prediction and generation. We prove that pervasive temperature choice of one is generically suboptimal under parallel sampling even for an exact denoiser, and that worse predictions can yield better samples. Through these results, we argue for a broader evaluation principle of treating the deployed generator as the object of comparison, benchmarking it against tuned baselines, and assessing quality and diversity jointly and with more human-aligned measures.
Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive. A growing class of training-free accelerators reduces this cost by reusing cached intermediate features or forecasting future ones. To control draft drift, these methods sometimes compute an exact block feature for verification. Yet the resulting exact feature is typically used only to measure discrepancy or guide a later decision and is then discarded. We find that this previously computed feature can instead be reused for correction. Forwarding it at the verification site resets the local draft residual and reduces downstream feature error. Based on this observation, we introduce FeatFix, a local exact-feature correction method for cached diffusion inference. FeatFix operates at a fixed sparse set of layer--timestep sites. At each selected site, it replaces the complete draft block output with the exact output computed from the same incoming state, avoiding token- or channel-level partial replacement and full-timestep recomputation. Experiments across four image and video backbones show that FeatFix consistently accelerates generation, achieving a speedup of up to 6.70× over Vanilla while maintaining competitive output quality.
Hanshuai Cui, Zhiqing Tang, Zhi Yao +3
School of Artificial Intelligence, Beijing Normal University, Beijing 100875, China · Institute of Artificial Intelligence and Future Networks, Beijing Normal University, Zhuhai 519087, China