cs.AISep 24, 2026

Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

Authors: Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yilan Wei, Yankai Zeng, Bojun Lin

Organizations: Independent · Northwestern University · Pinterest, Inc.

Abstract

Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of MM tasks with LL binary paths per task under the hard budget (M+t)K(M+t)K, where each path costs at most KK responses or episodes. For fixed L≥3L \ge 3 and 0<α≤1/120 < α\le 1/12, the optimal expected width on the worst pure cohort is Θα,L([M(t+1)]−1/2)Θ_{α,L}([M(t+1)]^{-1/2}) when every task is observed and Θα,L([M(t+M)]−1/2)Θ_{α,L}([M(t+\sqrt{M})]^{-1/2}) when omission is allowed. The lower bounds cover adaptive hard-budget policies, and fixed random-subset designs attain both rates through disagreement certificates. A joint mean/disagreement interval turns the task-covering law into practical finite-budget inference. In an equal-budget LiveCodeBench replay with 16 models, 880 tasks, and five outputs per task, the task-covering design reduces median point-estimation MSE by 87.0% relative to pooled uniform sampling, while the Joint certificate produces narrower confidence intervals in 15/16 panels and reduces median interval width by 30.6%. Finite-regime analyses identify task coverage as the effective choice at the evaluated scale and characterize how cohort size and within-task agreement determine the useful operating region. Together, the sharp laws and fixed-budget evidence make replication and task coverage explicit design variables for information-efficient repeated evaluation.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Risk-Aware Adaptive Evaluation: Finding High-Impact Failures Under Limited Budgets

    Sep 30, 2026Priyanath Maji, Spandan Ghose ChowdhuryThompson SamplingToken Budget Allocation

  2. Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves

    Sep 30, 2026Sohail, Sarkar, Shakuntala BaichooModel AuditingTest-Time Scaling