Reusing the scores that select a Best-of-N winner can overstate its expected reward. We study evaluation from a fixed matrix of K independent scores per candidate for a policy that selects using J fresh scores. A single estimator based only on this matrix is exactly unbiased for expected judge reward under every independent, stable collection of candidate-specific score laws if and only if J<K, for every pool size M≥N≥2. At J=K−1, the selector deepens as K grows. For independent Gaussian scores with common variance and fixed M≥N≥2, the unbiased minimax risk in this regime is of order σ2/K, attained by Holdout; allowing bias improves the rate to σ2/K. For two candidates, we derive the minimum-variance unbiased estimator at known variance and the sharp asymptotic unbiased minimax constant 1/(π2), which Holdout attains without knowing the variance. The cyclic average over subsets and ties can be computed in O(MKlogM) operations. At fixed selector depth, cyclic evaluation of bounded scores has O(K−1) risk uniformly in pool size. The impossibility result concerns the fixed matrix: one additional fresh winner score permits unbiased evaluation of the all-K policy.
Figures & tables
Figure 1: Selection and evaluation use disjoint score columns. Selection uses the mean of J columns; evaluation averages the winner’s remaining K−J scores. The estimator averages over cyclic splits, all size- N subsets, and uniform ties. Holdout uses J=K−1 .
Figure 2: Exact MSE divided by σ2 for two independent Gaussian candidates at equal means. Holdout uses no variance input; the other unbiased curves assume known variance. The J=1 and J=K−1 policies have the same value here.
RMSE against Θ(1)
cFVF/(cHVH)
K
Cyclic split ( H )
Fixed split ( F )
Score reuse
4
0.240
0.379
0.845
0.824
8
0.138
0.271
0.831
0.793
Table 1: Full-matrix cyclic and adaptive fixed-split evaluation of the one-score target Θ(1) . The last column compares population variance at equal score-call budget, excluding generation and prompt overhead; values below one favor the fixed split.
M
N
Means
Holdout H
Plug-in P
Δ (MCSE)
4
2
All tied
0.0366
0.0391
−0.00255(0.00026)
4
2
Two tied best
0.0262
0.0244
+0.00180(0.00006)
4
2
Separated
0.0247
0.0243
+0.00034(0.00002)
8
4
All tied
0.0452
0.0798
−0.03453(0.00039)
8
4
Two tied best
0.0250
0.0223
+0.00265(0.00008)
8
4
Separated
0.0226
0.0222
+0.00032(0.00002)
Table 2: MSE for Θ(15) with fixed means, K=16 , and σ=1 . Parentheses give the Monte Carlo standard error of the paired difference Δ=MSE(H)−MSE(P) ; positive Δ favors P .
ρ
Evaluation
Selector law
Total bias
0.0
−0.0001(0.0003)
+0.0000(0.0000)
−0.0001(0.0003)
0.1
+0.0551(0.0004)
−0.0160(0.0001)
+0.0391(0.0004)
0.4
+0.1832(0.0005)
−0.0478(0.0002)
+0.1353(0.0005)
0.8
+0.3074(0.0006)
−0.0720(0.0003)
+0.2355(0.0006)
Table 3: Bias against the independent-score target at each ρ . Entries are paired mean differences over 40,000 pools, with Monte Carlo standard errors in parentheses.