Diffusion models are a powerful generative paradigm used across multimedia and scientific applications. Guided diffusion methods impose requirements on the generation by adding the gradient of a differentiable loss (the guidance function) as a drift term during inference. The weight of this drift (the guidance scale) is critical for the trade-off between data quality and requirement satisfaction. To achieve both of these goals, guided diffusion must resort to small guidance scales and lengthy sampling, incurring high computational costs. This work proposes learned end-to-end guidance schedules (LEEGS) to achieve these objectives with fewer sampling steps. LEEGS trains a time-dependent schedule by minimizing the guidance function over a small set of examples using stochastic gradient descent. Backpropagating through guided sampling is computationally expensive, so LEEGS uses an approximation of the gradient that cuts training time by a factor of 4. We evaluate LEEGS on diverse guidance tasks, including (a) image inpainting, (b) noisy image inverse problems, (c) face-ID-guided generation, and (d) forward and inverse PDE problems, outperforming baselines at equal budget (50 or 100 NFEs), or matching constant guidance with only 10% of the steps.
Figures & tables
Figure 1
Dataset
Method
NFEs
Wide
Narrow
S. res. (2x)
Alternate Lines
Half
Imagenet
RePaint
200
0.186
0.18
0.622
0.394
0.368
LEEGS (ours)
200
0.182
0.151
0.147
0.111
0.346
RePaint
≈ 2400
0.134
0.064
0.183
0.089
0.304
CelebA-HQ
RePaint
200
0.128
0.128
0.205
0.049
0.283
LEEGS (ours)
200
0.102
0.080
0.065
0.042
0.229
RePaint
≈ 2400
0.059
0.028
0.029
0.009
0.165
Table 3: Comparison for image inpainting tasks for CelebA-HQ and Imagenet. Green (bold): best (non-reference rows).
Figure 3
Method
NFEs
Poisson
Helmholtz
Darcy
Forward
Inverse
Forward
Inverse
Forward
Inverse
Const.
50x2
19.34%
33.11%
27.56%
32.15%
21.54%
15.58%
Linear increase
50x1
27.59%
46.43%
39.72%
45.42%
24.70%
20.98%
Linear decrease
50x1
10.25%
29.81%
17.07%
29.84%
11.03%
14.05%
LEEGS (ours)
50x1
5.80%
29.36%
10.03%
29.71%
11.33%
17.84%
Const.
100x2
15.62%
30.10%
22.96%
29.54%
18.17%
16.71%
Table 5: Comparison for forward and inverse PDE problems with sparse observations. Green (bold): best; blue : second best (non-reference rows).
Figure 3: Learned guidance schedules for the forward Helmholtz problem (left) and the noisy box inpainting problem (right). The schedules are similar across different step budgets for a particular problem, but differ between problems, highlighting the need for learning methods.
Appendix figures & tables72 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Ground truth and the 5 noisy measurement operators used for Tables 1 and 2 (FFHQ example).
Subtask
Avg. NFEs ( T=25 )
Avg. NFEs ( T=50 )
FFHQ
Box inpainting
55.7
102.8
Random inpainting
50.7
93.7
Super-resolution ( 4× )
54.5
104.7
Gaussian deblur
56.0
104.4
Motion deblur
65.8
114.2
Appendix
Table 6: Average number of NFEs used by CDIM per subtask at its two settings, T=25 and T=50 (nominally 50 and 100 NFEs), computed over the full test set of Tables 1 and 2 .
Subtask
Learning rate η
Epochs
Weight decay
FFHQ
Super-resolution ( 4× )
0.3
25 ( S=50 ) / 15 ( S=100 )
1e-4 ( S=50 ) / 5e-4 ( S=100 )
Box inpainting
0.2
25 ( S=50 ) / 25 ( S=100 )
0.0
Random inpainting
0.8
25 ( S=50 ) / 20 ( S=100 )
0.0
Gaussian deblur
0.5
25 ( S=50 ) / 25 ( S=100 )
0.0
Motion deblur
0.3
30 ( S=50 ) / 25 ( S=100 )
0.0
Appendix
Table 7: LEEGS training configuration for the noisy inverse problems (Tables 1 , 2 ).
Figure 5: Ground truth and the 5 noiseless inpainting mask operators used for Table 3 (CelebA-HQ example).
Mask
Learning rate η
Epochs
Weight decay
ImageNet
Wide
0.2
5
0.0
Narrow
0.2
5
0.0
Super-resolve ( 2× )
0.2
5
0.0
Alternating lines
0.2
5
0.0
Half
0.2
5
0.0
Appendix
Table 8: LEEGS training configuration for image inpainting (Table 3 ).
NFEs
Learning rate η
Epochs
Weight decay
50
0.1
15
0.0
100
0.1
10
0.0
Appendix
Table 9: LEEGS training configuration for Face-ID guidance (Table 4 ).
Problem
Learning rate η
Epochs
Weight decay
Poisson, forward
15.0
15 ( S=50 ) / 15 ( S=100 )
0.0
Poisson, inverse
500.0
15 ( S=50 ) / 15 ( S=100 )
0.0
Helmholtz, forward
15.0 ( S=50 ) / 10.0 ( S=100 )
15 ( S=50 ) / 15 ( S=100 )
0.0
Helmholtz, inverse
1500.0
15 ( S=50 ) / 10 ( S=100 )
0.0
Darcy, forward
30.0
15 ( S=50 ) / 15 ( S=100 )
0.0
Darcy, inverse
150000.0 ( S=50 ) / 100000.0 ( S=100 )
15 ( S=50 ) / 15 ( S=100 )
0.0
Appendix
Table 10: LEEGS training configuration for the PDE problems (Table 5 ).
Results table
S
Epochs
Wall-clock
Time/epoch
FFHQ (Table 1 )
50
25 – 30
1h26m – 1h57m
3.4 – 3.9 min
100
15 – 25
1h52m – 3h06m
7.5 min
ImageNet (Table 2 )
50
20
3h11m
9.6 min
100
15
4h44m
19.0 min
Inpainting (Table 3 )
200 (CelebA-HQ)
5
3h13m
38.7 min
200 (ImageNet)
5
3h14m
38.7 min
Appendix
Table 11: Training wall-clock time by results table and NFE budget S , on a single A100 GPU with no validation logging. Where tasks at a fixed S use different epoch counts, both the minimum and maximum are reported (as a range); a single value indicates every task at that S trains for the same number of epochs.
Figure 6: Held-out test-set reconstruction quality as a function of training dataset size n . Scatter points are the 10 individual dataset draws, and the vertical bar spans min to max, with the mean indicated by a red square. Left: FFHQ motion deblur. Right: Helmholtz forward.
Figure 7: Per-draw L1 distance of the learned guidance schedule from the draw-averaged schedule (all S=10 steps), as a function of training dataset size n . Left: FFHQ motion deblur. Right: Helmholtz forward.
Figure 8: Draw-averaged guidance schedule (mean over 10 dataset draws) as a function of diffusion step, compared across training dataset sizes n Left: FFHQ motion deblur. Right: Helmholtz forward.
Figure 9: Learned guidance schedules for LEEGS and the true gradient (FFHQ and ImageNet motion deblur, S=10 ). Each panel shows the 10 per-draw schedules (thin grey) and their mean (bold); the true gradient’s higher spread on ImageNet is visible in the wider band around the mean, especially at later diffusion steps.
Dataset
Mode
LPIPS ( ↓ )
L2 dist. from mean
L1 dist. from mean
FFHQ
LEEGS
0.341±0.004
2.05
4.50
True gradient
0.342±0.006
2.34
5.52
ImageNet
LEEGS
0.560±0.007
2.28
5.13
True gradient
0.594±0.041
7.03
17.89
Appendix
Table 12: LEEGS vs. the true gradient on motion deblur ( S=10 ; mean ± std over 10 draws). L1 / L2 dist. from mean is each draw’s schedule distance to the 10 -draw mean. On ImageNet the true gradient is both worse and far less stable than LEEGS (std 4 – 6× higher on LPIPS, distances ∼3× higher); on FFHQ the two are comparable.
Figure 10: Wall-clock cost per training iteration vs. inference budget S , for LEEGS and the true gradient, relative to plain sampling. Left: FFHQ box inpainting – true gradient ≈20× sampling, LEEGS ≈4× . Right: Helmholtz forward – true gradient ≈16× , LEEGS ≈4× .
Figure 11: Held-out LPIPS as a function of weight decay γ , one curve per S (discrete colorbar). Left: superresolution. Right: box inpainting.
Figure 12: Learned guidance schedules for the five FFHQ noisy inverse tasks against normalized diffusion time ( 0 : noisiest, 1 : clean). Left: raw. Right: each schedule divided by its mean. Top: S=50 . Bottom: S=100 . The schedules differ strongly in magnitude and, to a lesser extent, in shape.
Figure 13: Held-out LPIPS ( 300 FFHQ images) when sampling with the schedule learned for the row task under the measurement operator of the column task. The best entry per column is in bold. Left: original schedules. Right: off-diagonal schedules rescaled to the mean guidance of the column task’s schedule. Top: S=50 . Bottom: S=100 . The matched schedule is best in almost every column before rescaling; after rescaling, some tasks continue to remain selective with respect to shape.
Figure 14: Mean guidance value of the trained schedules as a function of S , with the linear fit used for rescaling. Left: FFHQ box inpainting. Right: Helmholtz forward. The mean guidance decreases monotonically with S .
Ssrc→Stgt
Trained
Interp.
Interp.+rescale
10→20
0.217
0.203
0.204
10→30
0.173
0.169
0.170
10→50
0.140
0.142
0.139
10→100
0.119
0.156
0.113
20→30
0.173
0.172
0.174
20→50
0.140
0.139
0.140
Appendix
Table 13: FFHQ box inpainting: LPIPS ( ↓ ) on 300 held-out samples. Bold : best of the three columns.
Ssrc→Stgt
Trained
Interp.
Interp.+rescale
10→20
11.70%
9.80%
10.01%
10→30
10.34%
9.13%
9.01%
10→50
9.33%
9.88%
8.73%
10→100
9.07%
15.05%
8.77%
20→30
10.34%
9.54%
9.77%
20→50
9.33%
9.09%
9.04%
Appendix
Table 14: Helmholtz forward: relative L2 error on u (%, ↓ ) on 300 held-out samples. Bold : best of the three columns.
Figure 15: Training LPIPS vs. epoch, full vs. terminal-only constraint set, one panel per diffusion step budget S .
Figure 16: Held-out LPIPS vs. S , full vs. terminal-only constraint set.
Subtask
Learning rate η
Epochs
Weight decay
Super-resolution ( 4× )
0.3
15
0.0
Box inpainting
0.2
30
0.0
Random inpainting
0.8
30
0.0
Gaussian deblur
0.5
25
0.0
Motion deblur
0.3
15
0.0
Appendix
Table 15: LEEGS-renoise training configuration, S=50 , K=5 .
Task
LEEGS
LEEGS-renoise
Relative change
Box inpainting
0.174
0.165
−5.2%
Random inpainting
0.196
0.181
−7.5%
Gaussian deblur
0.243
0.246
+1.3%
Motion deblur
0.212
0.227
+7.0%
Superresolution
0.222
0.224
+0.6%
Appendix
Table 16: LPIPS on the FFHQ test set ( 950 images), S=50 .
Figure 17: Learned LEEGS-renoise guidance schedules for the five FFHQ tasks ( S=50 , K=5 ) against normalized diffusion time ( 0 : noisiest, 1 : clean); color indicates the outer step. Every outer step except the last has a hill shape.
Subtask
Learning rate η
Epochs
Weight decay
Super-resolution ( 4× )
0.3
30
0.0
Box inpainting
0.2
20
0.0
Random inpainting
0.8
30
0.0
Gaussian deblur
0.3
20
0.0
Motion deblur
0.3
20
0.0
Appendix
Table 17: LEEGS training configuration for the S=10 FFHQ extension (Table 18 ).
Figure 18: Box inpainting, FFHQ, 10 NFEs.
Figure 19: Random inpainting, FFHQ, 10 NFEs.
Figure 20: Superresolution ( ×4 ), FFHQ, 10 NFEs.
Figure 21: Gaussian deblur, FFHQ, 10 NFEs.
Figure 22: Motion deblur, FFHQ, 10 NFEs.
Dataset
NFEs
Wide
Narrow
S. res. (2x)
Alternate Lines
Half
Imagenet
200
0.0019
0.0040
0.0016
0.0018
0.0041
CelebA-HQ
200
0.0010
0.0001
0.0001
0.0005
0.0016
Appendix
Table 21: Standard deviation across 3 draws for the LEEGS rows of Table 3 (image inpainting).
NFEs
FaceID Score
FID
50
0.0036
0.90
100
0.0119
0.85
Appendix
Table 22: Standard deviation across 3 draws for the LEEGS rows of Table 4 (Face-ID guidance).
NFEs
Poisson
Helmholtz
Darcy
Forward
Inverse
Forward
Inverse
Forward
Inverse
50x1
0.01%
0.06%
0.08%
0.25%
4.27%
1.54%
100x1
0.01%
0.19%
0.06%
0.18%
1.45%
5.28%
Appendix
Table 23: Standard deviation across 3 draws for the LEEGS rows of Table 5 (PDE problems).
Dataset
Method
NFEs
Wide
Narrow
S. res. (2x)
Alternate Lines
Half
Imagenet
LEEGS (ours)
200
0.0058
0.0045
0.0029
0.0020
0.0030
RePaint
200
0.0093
0.0075
0.0021
0.0079
0.0014
CelebA-HQ
LEEGS (ours)
200
0.0022
0.0011
0.0003
0.0002
0.0013
RePaint
200
0.0053
0.0045
0.0023
0.0002
0.0017
Appendix
Table 30: Variance of LPIPS across test samples ( ↓ ) for image inpainting. Green (bold): best.
Dataset
Method
NFEs
Wide
Narrow
S. res. (2x)
Alternate Lines
Half
Imagenet
LEEGS (ours)
200
20.28
24.20
27.76
29.74
13.71
RePaint
200
18.88
21.69
16.81
21.65
13.47
CelebA-HQ
LEEGS (ours)
200
24.99
29.00
32.76
35.00
16.41
RePaint
200
21.40
24.73
29.24
36.46
13.88
Appendix
Table 31: PSNR in dB ( ↑ ) for image inpainting. Green (bold): best.
Dataset
Method
NFEs
Wide
Narrow
S. res. (2x)
Alternate Lines
Half
Imagenet
LEEGS (ours)
200
0.777
0.808
0.834
0.880
0.594
RePaint
200
0.784
0.794
0.295
0.635
0.584
CelebA-HQ
LEEGS (ours)
200
0.879
0.917
0.941
0.964
0.720
RePaint
200
0.847
0.868
0.845
0.966
0.670
Appendix
Table 32: SSIM ( ↑ ) for image inpainting. Green (bold): best.
Diffusion models have achieved strong performance in image, text-to-image, and video generation, where conditional generation is often controlled by classifier-free guidance (CFG). CFG improves condition consistency by increasing a guidance weight, but stronger guidance typically reduces diversity and distributional coverage. It remains unclear how this consistency-coverage trade-off should be controlled across the reverse trajectory, since the distribution induced by CFG is not simply the fixed-time tilted distribution given by the guided score field. To address this issue, we propose an information-theoretic framework for CFG schedule optimization. Our approach uses a clean endpoint reference to specify the desired consistency-coverage trade-off, while optimizing the actual distribution induced by the guided sampler toward this reference. We derive trajectory-level formulas to estimate the objective from samples and score evaluations, avoiding explicit density estimation. On ImageNet-512 with EDM-XXL and COCO with SD-XL, the learned schedules achieve competitive or improved trade-offs over constant guidance and allocate guidance selectively across noise levels.
Haobo Chen, Xiangxiang Xu, Yuheng Bu
Department of Computer Science University of California, Santa Barbara · Department of Computer Science University of Rochester
Recently, diffusion models have been widely adopted in generative modeling and have served as foundational models for many image generation tasks. To control the generation without costly re-training or fine-tuning, many works seek inference-time guidance methods to steer the latent via a differentiable objective at inference time. However, these methods cannot effectively preserve the original Gaussian distribution because they introduce distributional drift, thereby degrading the sample quality. To address this gap, we propose DiffRGD, a distribution-aware guidance framework that explicitly preserves the latent Gaussian structure. DiffRGD formulates each sampling step as a constrained optimization problem on a spherical manifold induced by the latent Gaussian distribution, and solves it efficiently via Riemannian Gradient Descent (RGD). DiffRGD is a plug-and-play method that can be seamlessly integrated into any pre-trained diffusion model. Extensive experiments demonstrate that DiffRGD outperforms previous methods in most image restoration and conditional generation tasks. Our project page is available at https://diffrgd.github.io/.
Jia-Wei Liao, Li-Xuan Peng, Mei-Heng Yueh +3
1National Taiwan University · 4Academia Sinica · 2National Tsing Hua University +1
Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then held fixed across prompts and denoising timesteps, even though different prompts and stages of generation can benefit from different parameter values. We introduce LeSAMP, a framework for learning prompt-conditioned, timestep-varying sampling parameters. We formulate parameter selection as a reinforcement learning problem: Given a user prompt, a large language model is trained to emit schedules for the chosen sampling parameters. We optimize our model using rewards from human preference models and VLM-as-a-judge. We evaluate our model on Flux.1 [dev] and Stable Diffusion 3.5, and find that compared to baselines, LeSAMP has a win rate of up to 68.12% using human preference scores and 73.37% using VLM-as-a-judge. These gains are validated in a user study where we achieve win rates of up to 59.46% over previous baselines. Our results suggest that learned sampling-parameter policies provide a complementary approach to existing post-training methods for improving diffusion model outputs.