Test-time compute can improve mathematical reasoning, but can short-budget runs predict how mathematical reasoning scales with additional compute? We introduce a Discovery--Execution (DE) framework that predicts the aggregate held-out scaling curves through a convolution of strategy discovery and conditional execution. From independent short-budget attempts and oracle-sketch-conditioned runs, the framework estimates cumulative success along held-out reasoning trajectories under alternate compute allocations. We evaluate four models on 35 fresh Olympiad problems and non-geometry problems from IMO-ProofBench Advanced. Under the DE framework, near-saturated execution predicts geometric scaling, as observed for the GPT models. For Claude Opus 4.8, incorporating measured execution substantially improves held-out forecasts over geometric extrapolation across one- and two-arm allocations. As a secondary application, regularized DE (R-DE) decisions to continue or restart yield lower average regret than the best model-specific retrospective policy. Together, these results show that measuring conditional execution provides information about longer reasoning that short-budget success rates do not always capture.
Figures & tables
Figure 1: Predicting the value of longer reasoning. Short budget runs measure success within one block; runs supplied with an oracle strategy sketch measure execution. Together, these measurements estimate discovery and execution, whose convolution predicts held-out unaided success across compute allocations. Each tile is one block (200k output tokens); N independent attempts receive K blocks each, with NK=8 .
Figure 2: Measured execution and held-out Parallel-1 scaling ( N=1 ). Top: sketch-conditioned and unaided successes, with dashed R-DE predictions. Bottom: difference between observed and predicted solved counts. Each condition has 171 trajectories across 57 problems (3 seeds each).
N=1
N=2
Model
SG
R-SG
SCT
DE
R-DE
SG
R-SG
SCT
DE
R-DE
Muse
8.26
8.31
11.52
6.54
6.25
5.61
5.58
10.49
6.26
5.58
GPT-5.4
3.96
3.98
20.46
4.25
3.93
2.83
2.83
9.97
2.32
2.16
GPT-5.5
2.40
1.86
13.80
3.04
3.66
3.59
3.03
6.47
3.88
3.46
Opus
7.16
7.20
11.81
3.68
2.87
9.16
9.10
5.48
6.70
2.96
Average
5.94
5.92
14.84
4.57
4.36
5.84
5.73
8.39
5.11
3.76
Table 1: Prediction RMSE (solved-trial counts) across 57 problems and 171 trials per model. Bold marks the lowest error within each row and allocation.
N
Execution
Trials
SG
DE
1
Saturated
87
2.40
2.65
Unsaturated
84
5.50
1.86
2
Saturated
87
2.17
2.50
Unsaturated
84
8.25
5.22
Table 2: Opus prediction RMSE (solved-trial counts), split by first-block execution saturation. Bold marks row minima. DE helps when execution is not saturated.
Model
Always continue
Always restart
Best fixed
R-DE
Muse Spark 1.2
9.74
0.57
0.57
1.57
GPT-5.4
3.14
3.62
3.14
2.29
GPT-5.5
2.10
1.19
1.19
0.76
Claude Opus 4.8
2.24
3.14
2.24
1.38
Average
4.30
2.13
1.79
1.50
Table 3: Mean next-block regret (solved trials lost), k=1,…,7 . Best fixed selects each model’s better constant action retrospectively. Bold marks row minima.
Figure 3: Next-block regret across inference budgets. Solved trials lost; lower is better. Budgets include the additional block.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: AOBench subtopics. Slice labels give problem counts; geometry is excluded from both benchmark cohorts.
Figure 5: Predicting two-arm success ( N=2 ). Top: observed successes and R-DE predictions. Bottom: observed minus predicted solved counts. Each curve represents 171 allocation trials across 57 problems; the horizontal axis gives total compute across both arms.
Model
SG
R-SG
SCT
DE
R-DE
Muse Spark 1.2
17.06
16.97
19.40
18.77
16.64
GPT-5.5
1.85
1.14
1.71
2.27
1.57
Appendix
Table 4: Prediction RMSE in solved-trial counts for N=4 , evaluated across both checkpoints of four arms with two blocks each. Each model has 57 problems and 171 allocation trials. Predictions use the unchanged intervention fit. Bold marks row minima.
Figure 6: Dataset-separated scaling. Solid lines show observed successes; dashed lines show R-DE predictions. Colors identify each dataset and their combined total.
Model
SG
R-SG
SCT
DE
R-DE
Muse Spark 1.2
1.27
1.23
1.26
1.19
1.09
GPT-5.4
1.22
1.21
1.51
1.20
1.16
GPT-5.5
0.71
0.69
1.10
0.74
0.71
Claude Opus 4.8
0.93
0.92
1.31
0.92
0.86
Average
1.06
1.04
1.31
1.03
0.97
Appendix
Table 5: Problem-level RMSE in solved-trajectory counts, using the six individual Parallel-2 trajectories per problem over blocks 1–4. Bold marks row minima.
Combined RMSE
RMSE reduction relative to baseline
Model
SG
R-SG
R-DE
SG minus R-DE
R-SG minus R-DE
Muse
7.06
7.08
5.93
1.13[−0.05,1.80]
1.15[−0.06,1.84]
GPT-5.4
3.44
3.45
3.17
0.27[−0.85,1.26]
0.28[−0.84,1.27]
GPT-5.5
3.05
2.52
3.56
−0.51[−0.77,0.18]
−1.04[−1.27,−0.02]
Opus
8.22
8.20
2.91
5.31[2.29,5.95]
5.29[2.29,5.92]
Appendix
Table 6: Combined held-out prediction accuracy across N=1 and N=2 . Equal-allocation RMSE in solved-trial counts. Reductions are baseline minus R-DE, so positive values favor R-DE; brackets give 95% paired target-bootstrap percentile intervals.
Table 8: Solved 1× trajectories on 38 matched problems (24 AOBench, 14 IMO-ProofBench), with three seeds each. Δ%=100(alternate−original)/original .
Model
Trajectories
γ
ΔlogL
Max. count change
Muse Spark 1.2
171
0.692
7.58
12.48
GPT-5.4
171
1.000
0.00
0.00
GPT-5.5
171
0.972
0.19
0.79
Claude Opus 4.8
171
0.974
0.12
0.86
Appendix
Table 9: Geometric-discovery sensitivity on all 57 problems. ΔlogL is the improvement over γ=1 ; the final column is the largest absolute change in predicted solved trajectories across the eight checkpoints.
Model
Trajectories
No sketch
Shuffled sketch
Oracle sketch
Muse Spark 1.2
105
24
13
48
GPT-5.4
105
41
38
81
GPT-5.5
105
83
77
104
Claude Opus 4.8
105
45
41
73
Appendix
Table 10: Sketch-content control on all 35 AOBench problems and three seeds. The last three columns count solved standalone 1× trajectories (score ≥5/7 ); seeds are counted individually.
Model
Parallel-8 (%)
Parallel-1, block 1 (%)
Difference (pp)
Muse Spark 1.2
27.0
32.7
+5.8
GPT-5.4
47.3
49.1
+1.8
GPT-5.5
80.6
80.7
+0.1
Claude Opus 4.8
51.6
53.2
+1.6
Appendix
Table 11: First-block success across all 57 problems: 1,368 Parallel-8 attempts and 171 Parallel-1 first-block outcomes per model. Differences are Parallel-1 minus Parallel-8, in percentage points (pp), computed before rounding.
Model
DE
Flat execution
Shared execution
Muse Spark 1.2
6.54
5.22
17.17
GPT-5.4
4.25
4.35
8.41
GPT-5.5
3.04
5.31
5.67
Claude Opus 4.8
3.68
19.47
21.20
Appendix
Table 12: Execution ablations: N=1 RMSE in solved-trajectory counts across eight checkpoints, with 171 held-out trajectories per model. Lower is better; bold marks row minima.
Figure 7: Muse Parallel-4 observed and predicted successes. Solved counts out of 171 trials for R-DE and the three post-hoc checks. Total budgets of 4× and 8× give one and two blocks per arm, respectively.
R-DE condition
RMSE
Final predicted / observed
R-DE
16.6
78.5 / 93
First-block success
1.7
94.4 / 93
Flat execution
17.9
75.9 / 93
Discovery decay
17.5
76.7 / 93
Appendix
Table 13: Muse N=4 diagnostics. RMSE uses both checkpoints in Figure 7 ; final counts are at two blocks per arm, out of 171 trials.
R-DE
R-SG
Model
All 57 + oracle
Train 38 + oracle
Train 38, no oracle
Train 38
Muse Spark 1.2
4.57±1.90
4.60±1.90
4.62±2.08
5.05±2.13
GPT-5.4
4.19±2.15
4.16±2.05
4.08±1.92
4.12±2.12
GPT-5.5
3.78±1.83
3.54±1.66
3.56±1.53
2.95±1.41
Claude Opus 4.8
3.62±1.47
3.40±1.46
5.04±2.67
6.24±2.09
Appendix
Table 14: Transfer across problems. Mean ± SD of solved-count RMSE over 50 splits, each using the six individual Parallel-2 trajectories for each of 19 test problems (114 trajectories), evaluated at blocks 1 – 4 with N=1 . All conditions except the first R-DE column fit priors on the 38 training problems.
Figure 8: Automated versus human proof scores for all 50 completed audits. Percentages are normalized within each row; n gives the number of proofs in that row.
Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency constraint, models must decide how to divide limited inference compute among them. We introduce an exam-style evaluation framework for studying this setting, in which a model must distribute one shared token budget across questions with different difficulty and point values to maximize its total score. Across several open and frontier reasoning models, we find that models fail to allocate a shared budget strategically across questions of varying difficulties and values. Models behave largely as greedy sequential solvers: they prioritize questions by presentation order, front-load effort on early questions, and remain insensitive to value, with these tendencies becoming more pronounced as the number of questions grows. Explicit planning prompts spread compute more evenly but do not produce value- or difficulty-aware prioritization. The same behavioral pattern extends from mathematical to code reasoning. These findings establish global budget allocation as a distinct capability that is not captured by conventional per-question evaluation and remains a challenge for current reasoning models.
Chenrui Fan, Yize Cheng, Ming Li +3
University of Maryland, College Park · MBZUAI, UAE
Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought reasoning. However, recent studies reveal that sequential test-time scaling often yields diminishing or even negative returns, as longer traces exhibit increased uncertainty, error compounding, and drift from the original problem. We propose ThinkRetrieve, a test-time scaling framework that augments the reasoning traces of LRMs with dynamically retrieved solved examples at each reasoning step. Given an external corpus of problems paired with step-by-step solutions, ThinkRetrieve retrieves relevant exemplars at each intermediate step and injects them directly into the thinking trace, providing the model with guidance on how to reason rather than merely what facts are relevant. Experiments across five reasoning models (1.5B--8B parameters) on GSM-8K, MATH-500, AIME 2025, and SciQ demonstrate that ThinkRetrieve consistently improves accuracy over standard test-time scaling, with relative gains of up to 60% on AIME 2025.
Large Reasoning Models (LRMs) achieve strong performance on mathematical reasoning tasks but remain unreliable on challenging instances. Existing test-time scaling methods, such as repeated sampling, self-correction, and tree search, improve performance at the cost of increased computation, yet often exhibit diminishing returns on hard problems. We observe that output disagreement is strongly correlated with instance difficulty and prediction correctness, providing a useful signal for guiding instance-level strategy selection at test time. Based on this insight, we propose a training-free framework that formulates test-time scaling as an instance-level routing problem, rather than allocating more computation within a single strategy, dynamically selecting among different scaling strategies based on output disagreement. The framework applies lightweight resolution for consistent cases, majority voting for moderate disagreement, and rewriting-based reformulation for highly ambiguous instances. Experiments on seven mathematical benchmarks and three models show that our method improves accuracy by 3% - 7% while reducing sampling cost compared to existing approaches.
Zhimin Lin, Yixin Ji, Jinpeng Li +5
School of Computer Science and Technology, Soochow University · Department of Foundation Model, 2012 Labs, Huawei · 3Harbin Institute of Technology, Shenzhen (HITSZ)