Efficient Best-of-N policy evaluation for inference-time alignment
Organizations: LMU Munich & MCML · Carnegie Mellon University · Meta (work in personal capacity)
Abstract
Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods. In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods. We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone. We then develop a doubly robust estimator of the BoN policy value (BoN-DR) that efficiently reuses a shared auxiliary sample pool across candidate budgets. We establish valid asymptotic inference even under reward estimator misspecification and prove the efficiency of our BoN-DR estimator. Since larger budgets can amplify errors in the score function and lead to reward overoptimization, we derive two selection rules: (i) maximizing the estimated policy value and (ii) maximizing a lower confidence bound on the improvement over the reference policy, which accounts for estimation uncertainty and provides a no-harm guarantee. Across synthetic experiments and GSM8K with multiple reference and reward models, our framework accurately estimates BoN policy values and selects effective sampling budgets.
Figures & tables
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Reward model | Estimator | Bias | RMSE | CI width | Coverage |
| OASST-RM-2.1, reference model Gemma-2-2B | |||||
| Plug-in | 0.097 | 0.097 | 0.002 | 0.000 | |
| BoN-IPW | 0.002 | 0.156 | 0.625 | 0.932 | |
| BoN-DR | 0.008 | 0.102 | 0.364 | 0.902 | |
| OASST-RM-2.1, reference model Llama-3.2-3B | |||||
| Plug-in | 0.029 | 0.029 | 0.001 | 0.000 | |
| Reward estimator | Estimator | Bias | RMSE | Coverage |
| Reference model Gemma-2-2B | ||||
| Embedding LightGBM | Plug-in | 0.040 | 0.040 | 0.000 |
| BoN-IPW | 0.008 | 0.167 | 0.880 | |
| BoN-DR | -0.015 | 0.093 | 0.970 | |
| TF–IDF logistic | Plug-in | 0.022 | 0.022 | 0.000 |
| BoN-IPW | 0.008 | 0.167 | 0.880 | |
| Synthetic exp (misspecified) | GSM8K: Qwen3.5-4B / OASST | |||||
| Selection method | Gain | Harm | Gain (pp) | Harm | ||
| Naive BoN | -0.388 | 1.00 | 256 | -1.50 | 1.00 | 256 |
| Plug-in | -0.380 | 1.00 | 256 | -0.28 | 1.00 | 4 |
| BoN-IPW, Value-max | 0.068 | 0.00 | 8 | -0.84 | 0.77 | 32 |
| BoN-IPW, Safe-improvement | 0.062 | 0.00 | 4 | 0.00 | 0.01 | 1 |
| BoN-DR, Value-max | 0.069 | 0.00 | 8 | -0.82 | 0.73 | 64 |
| Naive BoN | Plug-in | BoN-IPW | BoN-DR | |||||
| Gain | Harm | Rule | Gain | Harm | Gain | Harm | Gain | Harm |
| OASST-RM-2.1, reference model Gemma-2-2B | ||||||||
| 4.23 | 0.00 | Value-max | 4.84 | 0.00 | 4.45 | 0.00 | 4.51 | 0.00 |
| Safe-improvement | 4.84 | 0.00 | 3.25 | 0.00 | 3.98 | 0.00 | ||
| OASST-RM-2.1, reference model Llama-3.2-3B | ||||||||
| -0.51 | 1.00 | Value-max | 0.33 | 0.00 | -0.03 | 0.44 | -0.09 | 0.53 |
| Naive BoN | Plug-in | BoN-IPW | BoN-DR | |||||
| Gain | Harm | Rule | Gain | Harm | Gain | Harm | Gain | Harm |
| Reward estimator error | ||||||||
| -0.388 | 1.00 | Value-max | 0.065 | 0.00 | 0.068 | 0.00 | 0.070 | 0.00 |
| Safe-improvement | 0.065 | 0.00 | 0.062 | 0.00 | 0.067 | 0.00 | ||
| Reward estimator error | ||||||||
| -0.388 | 1.00 | Value-max | 0.071 | 0.00 | 0.068 | 0.00 | 0.070 | 0.00 |