cs.CLOct 6, 2026

Holdout Best-of-N: Unbiased Evaluation and Its Cost

Authors: Shrey Shah, Yinheng Li

Abstract

Reusing the scores that select a Best-of-NN winner can overstate its expected reward. We study evaluation from a fixed matrix of KK independent scores per candidate for a policy that selects using JJ fresh scores. A single estimator based only on this matrix is exactly unbiased for expected judge reward under every independent, stable collection of candidate-specific score laws if and only if J<KJ<K, for every pool size M≥N≥2M\ge N\ge2. At J=K−1J=K-1, the selector deepens as KK grows. For independent Gaussian scores with common variance and fixed M≥N≥2M\ge N\ge2, the unbiased minimax risk in this regime is of order σ2/Kσ^2/\sqrt K, attained by Holdout; allowing bias improves the rate to σ2/Kσ^2/K. For two candidates, we derive the minimum-variance unbiased estimator at known variance and the sharp asymptotic unbiased minimax constant 1/(π2)1/(π\sqrt2), which Holdout attains without knowing the variance. The cyclic average over subsets and ties can be computed in O(MKlog⁡M)O(MK\log M) operations. At fixed selector depth, cyclic evaluation of bounded scores has O(K−1)O(K^{-1}) risk uniformly in pool size. The impossibility result concerns the fixed matrix: one additional fresh winner score permits unbiased evaluation of the all-KK policy.

Figures & tables

Explore similar work

CardsList
  1. Efficient Best-of-N policy evaluation for inference-time alignment

    Oct 7, 2026Jonas Schweisthal, Yuxin Wang, Athiya Deviyani +2

  2. Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

    Aug 8, 2026Fariya Afrin, Ibne Farabi ShihabEnsembleCovariance

  3. Instance-Optimal Estimation with Multiple LLM Judges on a Budget

    May 22, 2026Junghyun Lee, Sanghwa Kim, Yassir Jedra +2Large Language Model JudgesLlm-As-A-Judge