Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is reported as a scaling curve: accuracy against the number k of sampled answers. The curve is cheap to draw and expensive to trust. A budget read off it is chosen after looking at every point, so only a band that covers all budgets at once protects the choice, and on a 100-question benchmark a fixed exact-binomial design needs 192,000 generated answers to certify 64 budgets to within ±1/32 at 95%. Most of that cost pays for the wrong uncertainty. A benchmark is a fixed list of questions; at budget 64, about three quarters of the variance of a selected answer's correctness lies between questions, and an audit that revisits every question need not pay for it. We derive the minimax cost of certifying the whole curve, up to logarithmic factors. It has three parts: calibrating the tail of the score distribution, telling the questions apart, and within-question noise summed along the curve. At a single benchmark the last part sharpens to the variance of one answer's influence under the best allocation of answers to questions, which every valid audit pays and an audit that learns the allocation attains, up to a logarithm, as the precision grows. A paired audit built on an exponential inequality for two independent draws at the same question needs no pilot. On 185 held-out score pools it uses 0.74 times the answers of the cheapest competing certified audit at 64 budgets and 0.53 times at 1,024, and on a newly generated MMLU-Pro study it certified the curve with 79,133 answers, within 0.6% of what a cost law fitted beforehand predicted from the study's within-question variance. The same paths certify pass@k and majority voting, and the bands extend to populations of questions and to answers that depend on earlier ones.
Figures & tables
Figure 1: Why a certified curve can be cheap. (a) Accuracy curves pk(x) of single questions spread from 0 to 1, while the benchmark average rises smoothly (MATH500, answers from Llama-3.1-8B-Instruct, one reward model; 60 of the 250 questions are drawn). (b) Median share of the variance of the selected answer’s correctness that lies between questions, with interquartile band, over the 185 held-out pools. (c) Generated answers needed for 100 MMLU-Pro questions and 64 budgets: an unbiased estimate of every point, a band from a fixed exact-binomial design, and a band from the paired audit of Section 6 , both at half-width 1/32 with 95% simultaneous coverage.
Figure 2: The coupling behind Lemma 3 . The highest score of the whole path lies in block A or in block B , so the winner at budget k is the winner of one of the two blocks. Exchanging the unshared parts of the blocks swaps their winners and leaves the law of the path unchanged.
Figure 3: The price at one benchmark. (a) Γ of Theorem 6 against Σw at K=64 for the 185 held-out pools and the MMLU-Pro study (law of the first 4,000 reference answers per question). The median of Σw/Γ is 12. (b) The allocation of answers to questions that attains Γ on the MMLU-Pro study, sorted by share.
Figure 4: One paired audit, traced round by round (the stored run on the MATH500 pool of Figure 1 a, 250 questions, K=64 ). (a) Answers generated per question visit in each pair of rounds; paths stop at the largest unresolved budget. (b) Half-width of the band at three budgets; a budget retires when it reaches 1/32 , and the final look is at round 18. (c) Answers used so far, against the fixed exact-binomial design and the cheapest of the competing certified audits on the same pool. The audit stops after 12 rounds with 127,500 answers, and its band covers the exact curve.
Figure 5: Theorem 13 : the leading ratio of the minimax errors without and with known score percentiles, as a function of the balance ClogK/(KT) between generated answers C and labels T .
solutions from
envelope
variance-optimal
fixed-width
primal–dual gap
Llama-3-8B
6,681
6,319
6,113
0.215%
Llama-3-70B
5,398
5,822
5,142
0.233%
Table 1: Sufficient label draws for a band of half-width 1/32 at level 95% over K=100 budgets with known percentiles, on two complete HumanEval+ pools scored by Llama-3.1-70B unit tests ( Ma et al.,, 2025 ; Liu et al.,, 2023 ) . The designs receive the same pool; answer generation is not charged. Last column: relative gap between the feasible fixed-width proposal and its dual lower bound.
Figure 6: First budget at which the exact curve of a held-out pool comes within 1/32 of its best value over budgets up to 1024.
K
answers
interval
pools cheaper
labels
64
0.74
[0.67, 0.88]
175/185
0.73
256
0.66
[0.59, 0.73]
175/185
0.68
1024
0.53
[0.47, 0.58]
173/185
0.61
Table 2: Paired audit on the 185 held-out pools, relative to the cheapest competing certified audit on each pool (median ratio of five-run means). Brackets: 95% bootstrap interval over the eight answer sets. Labels are compared with the cheapest competitor in labels.
Figure 7: (a) Distribution of the ratio of generated answers, paired audit to the cheapest competing certified audit, over the 185 held-out pools. Left of the dashed line the paired audit is cheaper than every competitor on that pool. (b) Median ratio for the designs of the ablation in Section 8.3 ; every design uses the same retirement of resolved budgets.
Figure 8: (a) Cost ratio at K=64 for every stored pool and three precisions, against Mε2 ; the start-up of two complete rounds decides the ratio at the right. (b) Median ratio on the held-out pools, whose questions have 100 stored answers, and on a pool with 4,000 answers per question built from the MMLU-Pro reference stream. (c) Σw as a fraction of its worst case K/4 in the same two settings (held-out: median).
Figure 9: The cost law. (a) N≈2.16KM+15.13K/ε+12.88Σw/ε2 , fitted to two paired audits per stored pool at K=64 and ε=1/32 ( R2=0.958 ), against measured cost. The star is the new MMLU-Pro study: the coefficients were fixed before any of its answers existed, and its Σw comes from the separate reference stream. (b) One law aKM+L(bK/ε+cΣw/ε2) for all horizons and precisions ( R2=0.995 ).
curve
audit
answers
labels
grader calls
cost ($)
best-of- k
paired
79,133
6,160
256
1.91
best-of- k
all-subsets
78,933
55,590
385
1.91
pass@ k
paired
60,467
5,943
285
1.46
pass@ k
all-subsets
49,067
49,067
385
1.18
majority voting
paired
62,267
58,999
356
1.50
majority voting
all-subsets
49,067
49,067
385
1.18
Table 3: MMLU-Pro study, M=100 , K=64 , ε=1/32 , δ=0.05 : means over three audit streams. Grader calls reuse one grade per distinct question–answer pair; dollars are the generation cost at the price we paid per answer. The designs below the rule certify best-of- k with fixed sample sizes.
Figure 10: Certified curves for 100 MMLU-Pro questions (first audit stream): simultaneous bands of half-width at most 1/32 at level 95% from the paired audit and the all-subsets audit, with reference curves estimated from 409,448 separately generated answers. Shaded in (b): budgets that the pass@ k bands rule out as a best budget.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
design
K=64
K=256
K=1024
exact binomial with completion
1.49 / 1.37
1.98 / 1.47
2.74 / 1.66
nested exact binomial
1.80 / 1.70
2.13 / 1.75
2.73 / 1.78
rank-based stopping
1.37 / 1.57
1.54 / 1.68
1.90 / 1.88
nested rank-based stopping
1.36 / 1.63
1.55 / 1.73
1.87 / 1.94
fixed Hoeffding
2.15 / 1.99
2.69 / 2.00
–
fixed exact binomial
1.65 / 1.53
2.12 / 1.60
–
Appendix
Table 4: Median ratio of each certified design’s cost to the paired audit’s on the 185 held-out pools (answers / correctness queries). The fixed designs and the record design were run at K=64 and 256.
design
K=64
K=256
K=1024
nested exact binomial, random questions
1.42
1.45
1.45
nested exact binomial, balanced rounds
1.45
1.45
1.47
per-row betting, balanced rounds
0.86
0.75
0.63
per-row betting, matched bet and cap
0.91
0.81
0.70
paired audit
0.74
0.66
0.53
balanced / random nested exact binomial
1.01
1.00
1.03
Appendix
Table 5: Ablation on the 185 held-out pools: median ratio to the cheapest competing certified audit (five-run means).