cs.LGSep 30, 2026

Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves

Authors: Sohail, Sarkar, Shakuntala Baichoo

Organizations: Neel · PMCC AI Lab, Peter Munk Cardiac Centre, University Health Network, Toronto, Ontario, Canada

Abstract

Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is reported as a scaling curve: accuracy against the number kk of sampled answers. The curve is cheap to draw and expensive to trust. A budget read off it is chosen after looking at every point, so only a band that covers all budgets at once protects the choice, and on a 100-question benchmark a fixed exact-binomial design needs 192,000 generated answers to certify 64 budgets to within ±1/32\pm1/32 at 95%. Most of that cost pays for the wrong uncertainty. A benchmark is a fixed list of questions; at budget 64, about three quarters of the variance of a selected answer's correctness lies between questions, and an audit that revisits every question need not pay for it. We derive the minimax cost of certifying the whole curve, up to logarithmic factors. It has three parts: calibrating the tail of the score distribution, telling the questions apart, and within-question noise summed along the curve. At a single benchmark the last part sharpens to the variance of one answer's influence under the best allocation of answers to questions, which every valid audit pays and an audit that learns the allocation attains, up to a logarithm, as the precision grows. A paired audit built on an exponential inequality for two independent draws at the same question needs no pilot. On 185 held-out score pools it uses 0.74 times the answers of the cheapest competing certified audit at 64 budgets and 0.53 times at 1,024, and on a newly generated MMLU-Pro study it certified the curve with 79,133 answers, within 0.6% of what a cost law fitted beforehand predicted from the study's within-question variance. The same paths certify pass@kk and majority voting, and the bands extend to populations of questions and to answers that depend on earlier ones.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Test-Time Scaling via Budgeted Multi-Attribute Verification

    Sep 28, 2026Bo Xue, Ji Cheng, Shen-Huan Lyu +2Token Budget AllocationLarge Language Model Responses

  2. Budget Boundary Effects in Test-Time Mathematical Reasoning

    Sep 30, 2026Guilin Zhang, Ziqi Tan, Wulan Guo +5Mathematical ReasoningTest Time

  3. When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

    Jun 27, 2026Yong Yi Bay, Kathleen A. YearickTest-Time ScalingReasoning Benchmark