Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-N (BoN) sampling, which draws N i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporated only at the final selection stage without influencing the reverse diffusion trajectory during sampling. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-N Guidance (BoNG), a novel method that integrates the principle of BoN sampling directly into the reverse diffusion process. BoNG performs online BoN selection over denoising particles and adjusts the reverse diffusion process to steer the particle population toward higher-reward regions during generation. Specifically, by introducing an asymmetric guidance interaction among denoising particles, BoNG uses the current BoN particle as a guidance signal to the rest of the particle population. This particle-level interaction reshapes the sampling process toward higher-reward regions, enabling BoNG to improve not only the final best sample beyond Vanilla BoN sampling, but also the average quality of generated samples. Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56% of the comparisons against SMC and Vanilla BoN sampling. BoNG also supports multi-output capability, achieving 1.3× ImageReward score of the latest sample-based guidance method with a 1.6× speedup. We release the code at https://github.com/aailab-kaist/BoNG.
Figures & tables
Figure 1: Best-of- N Guidance (BoNG) (Left) Vanilla BoN sampling selects the highest-reward sample only after generation. (Middle) BoNG performs online BoN selection ( Algorithm 1 ) and uses the selected sample to guide the remaining particles ( Algorithm 2 ). (Right) BoNG consistently improves single-output performance over Vanilla BoN and SMC across various sample budgets.
Figure 2: Dynamics of BoNG . (Left) Algorithm Summary. (Right) Decomposition of BoN target score in Eq. 7 .
Figure 3: Qualitative comparison of generated samples from Vanilla, SMC, and BoNG (SD v1.5 DDPM 100 step sampler). Red boxes denote Best-of- N samples for each method. Vanilla sampling often fails to generate images which align well with specific prompts such as “below a vase” and “a blue potted plant”. SMC generates better aligned images, but resampling leads to premature particle collapse. In contrast, BoNG generates samples that remain well aligned with the prompt while preserving diverse layouts and appearances.
Performance (↑)
Backbone
Method
N=2
N=3
N=4
IR
CLIP
GenEval
IR
CLIP
GenEval
IR
CLIP
GenEval
SD v1.5 (0.9B)
Vanilla
0.317
0.277
0.473
0.545
0.281
0.496
0.682
0.283
0.520
(w/ DDIM 50 steps)
SMC
0.317
0.277
0.473
0.532
0.281
0.512
0.636
0.282
0.509
BoNG
0.407
0.280
0.480
0.659
0.283
0.507
0.793
0.285
0.535
SD v1.5 (0.9B)
Vanilla
0.396
0.279
0.481
0.605
0.282
0.513
0.742
0.284
0.525
Table 1: Best-of- N (BoN) performance on GenEval prompts with different numbers of generated images N , using SD v1.5 Rombach et al. (2022) and SDXL Podell et al. (2024) backbones. ImageReward (IR) is used as the reward for all methods. Bold entries denote the top-performing results within each backbone group.
Method
Performance (↑)
Cost (↓)
IR
CLIP
GenEval
Time (s)
Mem. (G)
Vanilla
-0.001
0.271
0.426
7.07
8.90
Vanilla (250 steps)
0.003
0.272
0.430
17.42
8.90
Vanilla Best-4-of-5
0.180
0.275
0.453
8.83
9.97
UG Bansal et al. (2024)
0.326
0.262
0.355
58.36
28.16
DATE Na et al. (2025)
0.364
0.274
0.438
32.89
24.71
Table 2: Average performance on GenEval prompts with SD v1.5 backbone. Metrics are averaged over the 4 generated images per prompt, and cost is measured for generating 4 images (5 for Best-4-of-5) on a single A100 GPU. Bold entries indicate the best results.
Figure 6
Mean (↑)
BoN (↑)
Method
HPSv2
IR
CLIP
GenEval
HPSv2
IR
CLIP
GenEval
Vanilla
0.263
−0.018
0.272
0.427
0.287
0.488
0.282
0.517
BoNG
0.276
0.316
0.277
0.487
0.288
0.495
0.282
0.514
Table 7
Figure 8
N=2
N=3
N=4
Method
IR
CLIP
GenEval
IR
CLIP
GenEval
IR
CLIP
GenEval
Vanilla
1.220
0.285
0.671
1.303
0.287
0.687
1.353
0.287
0.696
BoNG
1.252
0.285
0.680
1.341
0.287
0.696
1.370
0.286
0.700
Table 9
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Window
IR (↑)
CLIP (↑)
GenEval (↑)
Vanilla
0.713
0.284
0.532
[20,24]
0.753
0.282
0.523
[40,44]
0.875
0.286
0.559
[60,64]
0.849
0.287
0.526
[80,84]
0.866
0.287
0.528
Appendix
Table 6: Ablation over the window position tstart (SD v1.5, DDPM 100 steps, N=4 , K=4 , s=6 ; BoN performance).
Window
K
mean IR (↑)
Diversity (↑)
Vanilla
–
−0.001
0.313
[40,44]
4
0.019
0.319
[40,50]
10
0.252
0.316
[40,55]
15
0.465
0.262
[40,60]
20
0.576
0.216
[40,65]
25
0.641
0.176
Appendix
Table 7: Controlled scan over the window size K (SD v1.5, DDPM 100 steps, N=4 , s=1 , tstart=40 ).
s
mean IR (↑)
Diversity (↑)
Vanilla
−0.001
0.313
1
0.019
0.319
2
0.224
0.315
3
0.490
0.255
4
0.473
0.245
6
0.506
0.224
Appendix
Table 8: Scan over the guidance scale s (SD v1.5, DDPM 100 steps, N=4 , window [40,44] , K=4 ).
Figure 8 : ImageReward [ 36 ] and CLIP score [ 7 ] trade-offs between UG, LiDAR, and BoNG. Across all methods, we report results under guidance scale s∈{1,2.5,5,7.5,10} .
Figure 9 : Changes in Best-of- N samples of BoNG as the guidance scale s∈{1,2,3,4,5,6,7,8,9,10} varies. The prompts from the top row to the bottom row are “ a photo of a pizza ”, “ a photo of a blue elephant ”, “ a photo of a cow left of a stop sign ”, “ a photo of a baseball glove right of a bear ”, and “ a photo of a yellow car and an orange toothbrush ”.
Method
Score-net
VAE decoder
Reward
Backward
SMC
–
N⋅S
N⋅S
–
UG / DATE
–
N⋅T
N⋅T
required
LiDAR
n⋅δ
n
n
–
BoNG
–
N⋅K
N⋅K
–
Appendix
Table 9: Additional cost over Vanilla sampling in hardware-independent units. D : VAE-decoder pass, R : reward pass; T : diffusion steps; S : resampling steps of SMC; n , δ : lookahead samples and score evaluations per lookahead of LiDAR.
Component
Time (ms)
Per particle (ms)
Time %
TFLOPs
TFLOPs / particle
UNet fwd + CFG
238.88
14.93
20.6%
25.705
1.607
VAE decode
522.98
32.69
45.1%
40.232
2.515
Postprocess (tensor → PIL)
167.94
10.50
14.5%
0.000
0.000
Reward (ImageReward)
230.81
14.43
19.9%
2.205
0.138
Appendix
Table 10: Wall-clock breakdown of one guided denoising step (SD v1.5, N=16 , mean of 5 runs).