cs.LGOct 1, 2026

How Much Can Language Models Gain from Test-Time Computation?

Authors: Bangji Yang, Jingyuan Li, Jiajun Fan, Yi Evie Zhang, Ruihan Guo, Hongba Ma, Neil He, Chumeng Liang, +3 more

Organizations: University of Illinois at Urbana-Champaign · University of Washington · Tsinghua University

Abstract

How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that measures the test-time potential of a model across competition mathematics, competitive programming, and agentic workflows. SELF-POT separates candidate coverage from final accuracy on static tasks, tracks correctness transitions under revision, and measures protocol completion alongside task success in agentic environments. Under a unified budget rule, it compares Direct inference with parallel sampling and self-revision under fixed multiples of the Direct budget, and charges every model call, including selection and critique, in dollars. This design supports two kinds of comparison: the gain a model obtains from additional inference, and a lower-cost model with additional inference against a stronger model. Across five low-cost reasoning models on 350 sealed tasks, with Claude Opus 5.5 Direct as the reference, the returns depend on the domain, the selection rule, and failure handling. When we replay the retained programming candidate pools, public-example selection raises correct submissions from 376 to 453 of 500 scheduled cells while saving 12-49% of logical API cost across models, and simply retaining an available candidate when judging fails recovers 61 submissions at unchanged cost. On identical mathematics pools, judging with fallback yields 186 correct submissions versus 182 for voting, while voting saves 12-21% of logical API cost. These controlled replays show how selection and failure handling change the gains realized from the same generated candidates, and they quantify the marginal value of a model judge.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scaling Evaluation-time Compute with Reasoning Models as Evaluators

    Mar 25, 2025Seungone Kim, Ian Wu, Jinu Lee +8Large Reasoning ModelsTest Time

  2. Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

    Aug 4, 2026Mohsen Hariri, Weicong Chen, Nahal Shahini +11Test-Time ScalingInference-Time

  3. Multi-Agent Reasoning Improves Compute Efficiency: Pareto-Optimal Test-Time Scaling

    May 2, 2026Florian Valentin Wunderlich, Lars Benedikt Kaesberg, Jan Philip Wahle +2Test-Time ScalingMulti-Agent Reasoning