cs.AIAug 12, 2026

Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation

Authors: Rodrigo Guedes de SouzaAlison R. Panisson

Organizations: Federal University of Santa Catarina (UFSC) Brazil

Abstract

Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels (64--4,096), evaluating four models on three reasoning benchmarks (56,476 inferences). We report four findings: (i) 3--19% of items exhibit non-monotone behavior (accuracy decreasing with more budget), even after controlling for truncation, and this phenomenon is model-specific (cross-model overlap: 6--14%). (ii) Model rankings reverse across budgets on all benchmarks (p<0.01p {<} 0.01, McNemar). (iii) Oracle analysis reveals model complementarity up to +27.8+27.8pp, most pronounced at constrained budgets. (iv) A budget-aware router captures 14.1% of the oracle gap cross-domain; budget features help within-domain (+1.6+1.6 to +5.7+5.7pp) but are domain-specific and hurt transfer (1.2-1.2pp). These results argue for budget-conditioned evaluation protocols.

Explore similar work

CardsList
  1. How Inference Compute Shapes Frontier LLM Evaluation

    Jun 16, 2026Jessica McFadyen, Ole Jorgensen, Harry Coppock +2

  2. The Capability Frontier: Benchmarks Miss 82% of Model Performance

    Jun 25, 2026Bradley Fowler, Ryan Smith, Daniel Thi Graviet +8FrontiersPareto Frontier