Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-N (BoN) sampling, which draws N i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporated only at the final selection stage without influencing the reverse diffusion trajectory during sampling. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-N Guidance (BoNG), a novel method that integrates the principle of BoN sampling directly into the reverse diffusion process. BoNG performs online BoN selection over denoising particles and adjusts the reverse diffusion process to steer the particle population toward higher-reward regions during generation. Specifically, by introducing an asymmetric guidance interaction among denoising particles, BoNG uses the current BoN particle as a guidance signal to the rest of the particle population. This particle-level interaction reshapes the sampling process toward higher-reward regions, enabling BoNG to improve not only the final best sample beyond Vanilla BoN sampling, but also the average quality of generated samples. Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56% of the comparisons against SMC and Vanilla BoN sampling. BoNG also supports multi-output capability, achieving 1.3× ImageReward score of the latest sample-based guidance method with a 1.6× speedup. We release the code at https://github.com/aailab-kaist/BoNG.
Figures & tables
Figure 1: Best-of- N Guidance (BoNG) (Left) Vanilla BoN sampling selects the highest-reward sample only after generation. (Middle) BoNG performs online BoN selection ( Algorithm 1 ) and uses the selected sample to guide the remaining particles ( Algorithm 2 ). (Right) BoNG consistently improves single-output performance over Vanilla BoN and SMC across various sample budgets.
Figure 2: Dynamics of BoNG . (Left) Algorithm Summary. (Right) Decomposition of BoN target score in Eq. 7 .
Figure 3: Qualitative comparison of generated samples from Vanilla, SMC, and BoNG (SD v1.5 DDPM 100 step sampler). Red boxes denote Best-of- N samples for each method. Vanilla sampling often fails to generate images which align well with specific prompts such as “below a vase” and “a blue potted plant”. SMC generates better aligned images, but resampling leads to premature particle collapse. In contrast, BoNG generates samples that remain well aligned with the prompt while preserving diverse layouts and appearances.
Performance (↑)
Backbone
Method
N=2
N=3
N=4
IR
CLIP
GenEval
IR
CLIP
GenEval
IR
CLIP
GenEval
SD v1.5 (0.9B)
Vanilla
0.317
0.277
0.473
0.545
0.281
0.496
0.682
0.283
0.520
(w/ DDIM 50 steps)
SMC
0.317
0.277
0.473
0.532
0.281
0.512
0.636
0.282
0.509
BoNG
0.407
0.280
0.480
0.659
0.283
0.507
0.793
0.285
0.535
SD v1.5 (0.9B)
Vanilla
0.396
0.279
0.481
0.605
0.282
0.513
0.742
0.284
0.525
Table 1: Best-of- N (BoN) performance on GenEval prompts with different numbers of generated images N , using SD v1.5 Rombach et al. (2022) and SDXL Podell et al. (2024) backbones. ImageReward (IR) is used as the reward for all methods. Bold entries denote the top-performing results within each backbone group.
Method
Performance (↑)
Cost (↓)
IR
CLIP
GenEval
Time (s)
Mem. (G)
Vanilla
-0.001
0.271
0.426
7.07
8.90
Vanilla (250 steps)
0.003
0.272
0.430
17.42
8.90
Vanilla Best-4-of-5
0.180
0.275
0.453
8.83
9.97
UG Bansal et al. (2024)
0.326
0.262
0.355
58.36
28.16
DATE Na et al. (2025)
0.364
0.274
0.438
32.89
24.71
Table 2: Average performance on GenEval prompts with SD v1.5 backbone. Metrics are averaged over the 4 generated images per prompt, and cost is measured for generating 4 images (5 for Best-4-of-5) on a single A100 GPU. Bold entries indicate the best results.
Figure 6
Mean (↑)
BoN (↑)
Method
HPSv2
IR
CLIP
GenEval
HPSv2
IR
CLIP
GenEval
Vanilla
0.263
−0.018
0.272
0.427
0.287
0.488
0.282
0.517
BoNG
0.276
0.316
0.277
0.487
0.288
0.495
0.282
0.514
Table 7
Figure 8
N=2
N=3
N=4
Method
IR
CLIP
GenEval
IR
CLIP
GenEval
IR
CLIP
GenEval
Vanilla
1.220
0.285
0.671
1.303
0.287
0.687
1.353
0.287
0.696
BoNG
1.252
0.285
0.680
1.341
0.287
0.696
1.370
0.286
0.700
Table 9
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Window
IR (↑)
CLIP (↑)
GenEval (↑)
Vanilla
0.713
0.284
0.532
[20,24]
0.753
0.282
0.523
[40,44]
0.875
0.286
0.559
[60,64]
0.849
0.287
0.526
[80,84]
0.866
0.287
0.528
Appendix
Table 6: Ablation over the window position tstart (SD v1.5, DDPM 100 steps, N=4 , K=4 , s=6 ; BoN performance).
Window
K
mean IR (↑)
Diversity (↑)
Vanilla
–
−0.001
0.313
[40,44]
4
0.019
0.319
[40,50]
10
0.252
0.316
[40,55]
15
0.465
0.262
[40,60]
20
0.576
0.216
[40,65]
25
0.641
0.176
Appendix
Table 7: Controlled scan over the window size K (SD v1.5, DDPM 100 steps, N=4 , s=1 , tstart=40 ).
s
mean IR (↑)
Diversity (↑)
Vanilla
−0.001
0.313
1
0.019
0.319
2
0.224
0.315
3
0.490
0.255
4
0.473
0.245
6
0.506
0.224
Appendix
Table 8: Scan over the guidance scale s (SD v1.5, DDPM 100 steps, N=4 , window [40,44] , K=4 ).
Figure 8 : ImageReward [ 36 ] and CLIP score [ 7 ] trade-offs between UG, LiDAR, and BoNG. Across all methods, we report results under guidance scale s∈{1,2.5,5,7.5,10} .
Figure 9 : Changes in Best-of- N samples of BoNG as the guidance scale s∈{1,2,3,4,5,6,7,8,9,10} varies. The prompts from the top row to the bottom row are “ a photo of a pizza ”, “ a photo of a blue elephant ”, “ a photo of a cow left of a stop sign ”, “ a photo of a baseball glove right of a bear ”, and “ a photo of a yellow car and an orange toothbrush ”.
Method
Score-net
VAE decoder
Reward
Backward
SMC
–
N⋅S
N⋅S
–
UG / DATE
–
N⋅T
N⋅T
required
LiDAR
n⋅δ
n
n
–
BoNG
–
N⋅K
N⋅K
–
Appendix
Table 9: Additional cost over Vanilla sampling in hardware-independent units. D : VAE-decoder pass, R : reward pass; T : diffusion steps; S : resampling steps of SMC; n , δ : lookahead samples and score evaluations per lookahead of LiDAR.
Component
Time (ms)
Per particle (ms)
Time %
TFLOPs
TFLOPs / particle
UNet fwd + CFG
238.88
14.93
20.6%
25.705
1.607
VAE decode
522.98
32.69
45.1%
40.232
2.515
Postprocess (tensor → PIL)
167.94
10.50
14.5%
0.000
0.000
Reward (ImageReward)
230.81
14.43
19.9%
2.205
0.138
Appendix
Table 10: Wall-clock breakdown of one guided denoising step (SD v1.5, N=16 , mean of 5 runs).
We introduce the Noise-Tilted Reverse Kernel (NTRK), a reward-guided diffusion sampler that injects reward gradients through the noise term, leaving the pretrained reverse kernel unchanged and requiring only a single sample per step. Reward-guided sampling at inference time has greatly expanded the versatility of pretrained diffusion models. Yet existing methods face a trade-off. Gradient-based guidance shifts the reverse mean, steering generation but pushing intermediate states outside the region that the model was trained on and degrading quality. Search-based methods preserve quality but gain no gradient signal. No prior method achieves both. NTRK resolves this by keeping the reverse mean fixed and biasing the noise term toward high reward. This is enabled by a whitening operator, the central mechanism behind NTRK, which converts reward gradients into noise-compatible perturbations without losing their guiding signal. Across various reward alignment tasks, NTRK outperforms recent state-of-the-art baselines without losing sample quality. Remarkably, on aesthetic generation, NTRK surpasses the reward of the best baseline at 500 NFEs using only 25 NFEs, a 20 times reduction in compute.
Jisung Hwang, Yunhong Min, Jaihoon Kim +2
1KAIST · †This work was done in part while the author was a visiting researcher at The University of Tokyo. · 2The University of Tokyo
Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward models, pre-trained for semantic alignment, struggle to capture the essential perceptual qualities-such as aesthetics, composition, and visual harmony. In this work, we argue that a model capable of high-fidelity generation must possess a profound understanding of these visual attributes. Based on this insight, we introduce the Diffusion-based Reward Model (DRM), a novel paradigm that use the pre-trained diffusion model as a powerful evaluative backbone. A key advantage of the DRM is its unique ability to assess not only the final image but also the noisy intermediate latents at any stage of the generative process. We leverage this step-wise evaluative capacity in two ways. First, we propose Step-wise GRPO, a reinforcement learning algorithm that provides dense, per-step rewards to resolve the imprecise credit assignment problem in GRPO algorithm, leading to more stable and effective alignment. Second, we introduce Step-wise Sampling, a novel inference strategy that employs the DRM as a dynamic guide to evaluate multiple generation paths at each step, steering the process towards higher-quality outcomes. Extensive experiments confirm that our approach significantly enhances the final quality of generated images. Code: https://github.com/jjaxonx/DRM.
Jaxon Zhang, Binxin Yang, Hubery Yin +2
Peking University · Work done during an internship at WeChat Vision, Tencent Inc. · WeChat Vision, Tencent Inc.
Optimizing the noise samples of diffusion and flow models is an increasingly popular approach to align these models to target rewards at inference time. However, we observe that these approaches are usually restricted to differentiable or cheap reward models, the formulation of the underlying pretrained generative model, or are memory/compute inefficient. We instead propose a simple trust-region based search algorithm (TRS) which treats the pre-trained generative and reward models as a black-box and only optimizes the source noise. Our approach achieves a good balance between global exploration and local exploitation, and is versatile and easily adaptable to various generative settings and reward models with minimal hyperparameter tuning. We evaluate TRS across text-to-image, molecule and protein design tasks, and obtain significantly improved output samples over the base generative models and other inference-time alignment approaches which optimize the source noise sample, or even the entire reverse-time sampling noise trajectories in the case of diffusion models. Our source code is publicly available.
Niklas Schweiger, Daniel Cremers, Karnik Ram
Technical University of Munich, Germany · Munich Center for Machine Learning, Germany