Test-time compute can improve mathematical reasoning, but can short-budget runs predict how mathematical reasoning scales with additional compute? We introduce a Discovery--Execution (DE) framework that predicts the aggregate held-out scaling curves through a convolution of strategy discovery and conditional execution. From independent short-budget attempts and oracle-sketch-conditioned runs, the framework estimates cumulative success along held-out reasoning trajectories under alternate compute allocations. We evaluate four models on 35 fresh Olympiad problems and non-geometry problems from IMO-ProofBench Advanced. Under the DE framework, near-saturated execution predicts geometric scaling, as observed for the GPT models. For Claude Opus 4.8, incorporating measured execution substantially improves held-out forecasts over geometric extrapolation across one- and two-arm allocations. As a secondary application, regularized DE (R-DE) decisions to continue or restart yield lower average regret than the best model-specific retrospective policy. Together, these results show that measuring conditional execution provides information about longer reasoning that short-budget success rates do not always capture.
Figures & tables
Figure 1: Predicting the value of longer reasoning. Short budget runs measure success within one block; runs supplied with an oracle strategy sketch measure execution. Together, these measurements estimate discovery and execution, whose convolution predicts held-out unaided success across compute allocations. Each tile is one block (200k output tokens); N independent attempts receive K blocks each, with NK=8 .
Figure 2: Measured execution and held-out Parallel-1 scaling ( N=1 ). Top: sketch-conditioned and unaided successes, with dashed R-DE predictions. Bottom: difference between observed and predicted solved counts. Each condition has 171 trajectories across 57 problems (3 seeds each).
N=1
N=2
Model
SG
R-SG
SCT
DE
R-DE
SG
R-SG
SCT
DE
R-DE
Muse
8.26
8.31
11.52
6.54
6.25
5.61
5.58
10.49
6.26
5.58
GPT-5.4
3.96
3.98
20.46
4.25
3.93
2.83
2.83
9.97
2.32
2.16
GPT-5.5
2.40
1.86
13.80
3.04
3.66
3.59
3.03
6.47
3.88
3.46
Opus
7.16
7.20
11.81
3.68
2.87
9.16
9.10
5.48
6.70
2.96
Average
5.94
5.92
14.84
4.57
4.36
5.84
5.73
8.39
5.11
3.76
Table 1: Prediction RMSE (solved-trial counts) across 57 problems and 171 trials per model. Bold marks the lowest error within each row and allocation.
N
Execution
Trials
SG
DE
1
Saturated
87
2.40
2.65
Unsaturated
84
5.50
1.86
2
Saturated
87
2.17
2.50
Unsaturated
84
8.25
5.22
Table 2: Opus prediction RMSE (solved-trial counts), split by first-block execution saturation. Bold marks row minima. DE helps when execution is not saturated.
Model
Always continue
Always restart
Best fixed
R-DE
Muse Spark 1.2
9.74
0.57
0.57
1.57
GPT-5.4
3.14
3.62
3.14
2.29
GPT-5.5
2.10
1.19
1.19
0.76
Claude Opus 4.8
2.24
3.14
2.24
1.38
Average
4.30
2.13
1.79
1.50
Table 3: Mean next-block regret (solved trials lost), k=1,…,7 . Best fixed selects each model’s better constant action retrospectively. Bold marks row minima.
Figure 3: Next-block regret across inference budgets. Solved trials lost; lower is better. Budgets include the additional block.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: AOBench subtopics. Slice labels give problem counts; geometry is excluded from both benchmark cohorts.
Figure 5: Predicting two-arm success ( N=2 ). Top: observed successes and R-DE predictions. Bottom: observed minus predicted solved counts. Each curve represents 171 allocation trials across 57 problems; the horizontal axis gives total compute across both arms.
Model
SG
R-SG
SCT
DE
R-DE
Muse Spark 1.2
17.06
16.97
19.40
18.77
16.64
GPT-5.5
1.85
1.14
1.71
2.27
1.57
Appendix
Table 4: Prediction RMSE in solved-trial counts for N=4 , evaluated across both checkpoints of four arms with two blocks each. Each model has 57 problems and 171 allocation trials. Predictions use the unchanged intervention fit. Bold marks row minima.
Figure 6: Dataset-separated scaling. Solid lines show observed successes; dashed lines show R-DE predictions. Colors identify each dataset and their combined total.
Model
SG
R-SG
SCT
DE
R-DE
Muse Spark 1.2
1.27
1.23
1.26
1.19
1.09
GPT-5.4
1.22
1.21
1.51
1.20
1.16
GPT-5.5
0.71
0.69
1.10
0.74
0.71
Claude Opus 4.8
0.93
0.92
1.31
0.92
0.86
Average
1.06
1.04
1.31
1.03
0.97
Appendix
Table 5: Problem-level RMSE in solved-trajectory counts, using the six individual Parallel-2 trajectories per problem over blocks 1–4. Bold marks row minima.
Combined RMSE
RMSE reduction relative to baseline
Model
SG
R-SG
R-DE
SG minus R-DE
R-SG minus R-DE
Muse
7.06
7.08
5.93
1.13[−0.05,1.80]
1.15[−0.06,1.84]
GPT-5.4
3.44
3.45
3.17
0.27[−0.85,1.26]
0.28[−0.84,1.27]
GPT-5.5
3.05
2.52
3.56
−0.51[−0.77,0.18]
−1.04[−1.27,−0.02]
Opus
8.22
8.20
2.91
5.31[2.29,5.95]
5.29[2.29,5.92]
Appendix
Table 6: Combined held-out prediction accuracy across N=1 and N=2 . Equal-allocation RMSE in solved-trial counts. Reductions are baseline minus R-DE, so positive values favor R-DE; brackets give 95% paired target-bootstrap percentile intervals.
Table 8: Solved 1× trajectories on 38 matched problems (24 AOBench, 14 IMO-ProofBench), with three seeds each. Δ%=100(alternate−original)/original .
Model
Trajectories
γ
ΔlogL
Max. count change
Muse Spark 1.2
171
0.692
7.58
12.48
GPT-5.4
171
1.000
0.00
0.00
GPT-5.5
171
0.972
0.19
0.79
Claude Opus 4.8
171
0.974
0.12
0.86
Appendix
Table 9: Geometric-discovery sensitivity on all 57 problems. ΔlogL is the improvement over γ=1 ; the final column is the largest absolute change in predicted solved trajectories across the eight checkpoints.
Model
Trajectories
No sketch
Shuffled sketch
Oracle sketch
Muse Spark 1.2
105
24
13
48
GPT-5.4
105
41
38
81
GPT-5.5
105
83
77
104
Claude Opus 4.8
105
45
41
73
Appendix
Table 10: Sketch-content control on all 35 AOBench problems and three seeds. The last three columns count solved standalone 1× trajectories (score ≥5/7 ); seeds are counted individually.
Model
Parallel-8 (%)
Parallel-1, block 1 (%)
Difference (pp)
Muse Spark 1.2
27.0
32.7
+5.8
GPT-5.4
47.3
49.1
+1.8
GPT-5.5
80.6
80.7
+0.1
Claude Opus 4.8
51.6
53.2
+1.6
Appendix
Table 11: First-block success across all 57 problems: 1,368 Parallel-8 attempts and 171 Parallel-1 first-block outcomes per model. Differences are Parallel-1 minus Parallel-8, in percentage points (pp), computed before rounding.
Model
DE
Flat execution
Shared execution
Muse Spark 1.2
6.54
5.22
17.17
GPT-5.4
4.25
4.35
8.41
GPT-5.5
3.04
5.31
5.67
Claude Opus 4.8
3.68
19.47
21.20
Appendix
Table 12: Execution ablations: N=1 RMSE in solved-trajectory counts across eight checkpoints, with 171 held-out trajectories per model. Lower is better; bold marks row minima.
Figure 7: Muse Parallel-4 observed and predicted successes. Solved counts out of 171 trials for R-DE and the three post-hoc checks. Total budgets of 4× and 8× give one and two blocks per arm, respectively.
R-DE condition
RMSE
Final predicted / observed
R-DE
16.6
78.5 / 93
First-block success
1.7
94.4 / 93
Flat execution
17.9
75.9 / 93
Discovery decay
17.5
76.7 / 93
Appendix
Table 13: Muse N=4 diagnostics. RMSE uses both checkpoints in Figure 7 ; final counts are at two blocks per arm, out of 171 trials.
R-DE
R-SG
Model
All 57 + oracle
Train 38 + oracle
Train 38, no oracle
Train 38
Muse Spark 1.2
4.57±1.90
4.60±1.90
4.62±2.08
5.05±2.13
GPT-5.4
4.19±2.15
4.16±2.05
4.08±1.92
4.12±2.12
GPT-5.5
3.78±1.83
3.54±1.66
3.56±1.53
2.95±1.41
Claude Opus 4.8
3.62±1.47
3.40±1.46
5.04±2.67
6.24±2.09
Appendix
Table 14: Transfer across problems. Mean ± SD of solved-count RMSE over 50 splits, each using the six individual Parallel-2 trajectories for each of 19 test problems (114 trajectories), evaluated at blocks 1 – 4 with N=1 . All conditions except the first R-DE column fit priors on the 38 training problems.
Figure 8: Automated versus human proof scores for all 50 completed audits. Percentages are normalized within each row; n gives the number of proofs in that row.
School of Computer Science and Technology, Soochow University · Department of Foundation Model, 2012 Labs, Huawei · 3Harbin Institute of Technology, Shenzhen (HITSZ)