Sampling multiple solutions spends computation on intermediate deductions and unfinished arguments as well as final answers. We introduce ReSolve, a training-free inference procedure that reuses this candidate reasoning through selective generative moderation. An answer-distribution controller invokes a model to examine existing derivations when candidates disagree or lack a parseable answer, then incorporates the generated solution into a bounded loop. Under Hybrid scoring on 130 competition-mathematics problems evaluated with two independently sampled candidate pools, ReSolve obtains 100 and 99 correct answers, compared with 91 and 92 for voting over the same four candidates, with no correct-to-incorrect changes relative to that vote in either pool. Eight-sample self-consistency obtains 94 and 96 correct answers while consuming substantially more tokens; ReSolve uses 46.3% and 47.2% fewer tokens in the two evaluations. A controlled ablation removes visible derivations while retaining answer keys, vote counts, and the per-state output-cap rule, reducing accuracy from 100 to 93 correct despite increasing computation. Selective and always-on Uniform moderation both solve 97 problems, while selectivity reduces moderation tokens by approximately 54% and total pipeline tokens by 6.2%. These results support candidate reasoning as reusable inference computation. They do not establish an accuracy advantage over additional sampling or a distinct benefit from specialized route instructions.
Figures & tables
Figure 1 : Selective reuse of candidate reasoning. Four sampled solutions provide both an answer distribution and visible derivations. Agreement permits a direct decision; otherwise, a generative moderator checks, repairs, or completes candidate reasoning. New solutions update the pool until acceptance or the call limit. Proposer and moderator share existing model weights, and reference answers are absent at inference.
State
Histogram condition
Intervention
No final answer
nC=0
Complete or repair unfinished attempts
Agreement
One valid answer group
Commit without a new call
Dominant answer
Top share >τ
Audit the majority and competing evidence
Two groups
Two groups, no dominant answer
Adjudicate the competing arguments
Fragmented
More than two, no dominant answer
Reconcile useful deductions
Table 1: Controller states, evaluated in the displayed order. Shares use the number of parseable answers. In particular, agreement can occur with fewer than K parseable candidates. We use τ=0.5 .
Method
F2 correct
F2 (%)
F3 correct
F3 (%)
SC@4
91
70.00
92
70.77
SC@8
94
72.31
96
73.85
Direct32
82
63.08
80
61.54
Uniform
97
74.62
99
76.15
ReSolve
100
76.92
99
76.15
Oracle@4
99
76.15
100
76.92
Table 2: Hybrid results on two independent candidate pools for the same 130 problems. ReSolve and Uniform access only the first four candidates. Oracle rows use evaluation reference answers. The evaluations are repeated observations of the same questions.
Figure 2 : Accuracy and token cost with Hybrid scoring. (a) F2 controls, including the answer-only intervention; the star marks ReSolve. (b) Whole-pipeline token cost for ReSolve and SC@8 on both independent candidate pools. Each method’s total includes the proposals it uses and all moderation input/output tokens.
F2 configuration
Correct / 130
Calls
Moderation (M)
ReSolve (Full)
100
51
0.753
Answer-only
93
65
1.276
Uniform loop, selective
97
50
0.706
Uniform loop, always-on
97
136
1.522
Uniform one-call, selective
97
44
0.619
Uniform one-call, always-on
97
130
1.435
Table 3: F2 controls under Hybrid. Moderation tokens include inputs and outputs; all rows share an additional 11.711M proposal tokens. Full versus Answer-only removes derivations under the same per-state output-cap rule. Selective versus always-on Uniform adds agreement checks, shared across the two Uniform variants.
F2 initial state
N
vs. SC@4
vs. SC@8
vs. Uniform
Agreement
86
0/0
0/0
0/0
Dominant answer
18
2/0
2/0
3/0
Two groups
8
1/0
0/0
0/0
Fragmented
12
3/0
2/1
1/1
No answer
6
3/0
3/0
0/0
Table 4: F2 fixes/breaks by initial state under Hybrid. Comparators isolate same-pool voting, additional proposals, and the moderation instruction. Small strata are descriptive.
Example: a quartic root (CMIMC25, cache index 7). Problem: Let P(x)=x4+20x3+29x2−666x+2025>0 for real x . A first-quadrant root has the form r=(a+bi+c+di)/2 , with integer a,b,c,d . Find a+b+c+d . Reference answer: 322.
Initial candidate pool / SC@4
ReSolve
All four attempts lack an extractable final answer, so answer voting has no valid answer to select. However, the fourth attempt already contains \color[rgb]{0,0,0}P(x)=(x^{2}+10x-36)^{2}+(x+27)^{2}. This useful intermediate identity remains visible in the moderator’s actual input.
Checks the decomposition and completes the complex quadratic: x2+(10−i)x−(36+27i)=0. The discriminant is 243+88i , giving r=2−10+i+243+88i. Thus −10+1+243+88=322 .
Extracted answers: [⊥,⊥,⊥,⊥]
One moderation call. Answer: 322 ✓
Table 5 : Completing candidate reasoning on CMIMC 2025. The initial pool has no extractable final answer, but contains an identity that supports a correct completion. All descriptions are edited summaries, not verbatim model quotations.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
AIME24
AIME25
CMIMC25
HMMT25
Total
SC@4
26/26
24/24
25/24
16/16
91/90
Oracle@4
27/27
25/25
27/26
20/20
99/98
SC@8
27/27
25/25
26/25
16/16
94/93
Oracle@8
27/27
25/25
28/27
20/20
100/99
ReSolve
27/27
25/25
27/27
21/21
100/100
Uniform
27/27
26/26
25/25
19/19
97/97
Appendix
Table 6: F2 per-task correct counts as Hybrid/CV. Task sizes are 30, 30, 40, and 30. Hybrid and CV score the same outputs; the main text uses Hybrid.
Pool
Score
Comparator
Fix/break
Exact p
Holm p
F2
Hybrid
SC@8
7/1
0.07031
0.14062
F2
Hybrid
Uniform
4/1
0.37500
0.37500
F2
Hybrid
SC@4
9/0
0.00391
—
F3
Hybrid
SC@8
7/4
0.54883
1.00000
F3
Hybrid
Uniform
2/2
1.00000
1.00000
F3
Hybrid
SC@4
7/0
0.01562
—
Appendix
Table 7: Within-pool tests for ReSolve. Holm correction covers SC@8 and Uniform within each pool and score. SC@4 is a secondary comparison. Combined problem-level tests use a different analysis described in the text.
Method
F2 rule
F2 CV
F2 Hybrid
F3 CV
F3 Hybrid
ReSolve
100
100
100
97
99
Uniform
97
97
97
97
99
SC@8
94
93
94
94
96
SC@4
91
90
91
91
92
Direct32
81
81
82
79
80
Appendix
Table 8: Scoring sensitivity on shared outputs. F2 rule values are from the audit; F2 Hybrid/CV values are reconciled to archived outcomes. Provisional manual judgments do not replace recorded scores.
CV contrast
Discordant problems
Clipped inputs
Final answers lost
ReSolve vs. SC@8
9
2
0
ReSolve vs. Uniform
5
0
0
ReSolve vs. Answer-only
8
2
0
ReSolve vs. Direct32
23
17
0
ReSolve vs. Oracle@4
6
17
0
ReSolve vs. Oracle@8
7
39
0
Appendix
Table 9: Reported replay of CV clipping on discordant F2 comparisons. Oracle comparisons inspect candidate scoring inputs as well as moderator outputs, so input counts can exceed problem counts. Rows overlap and must not be summed. No existing final answer is removed; clipped Answer-only outputs already lacked final answers before clipping.
Initial state
N
ReSolve
SC@4
SC@8
Uniform
Oracle@4
Agreement
86
74
74
74
74
74
Dominant answer
18
14
12
12
11
14
Two groups
8
5
4
5
5
5
Fragmented
12
4
1
3
4
6
No answer
6
3
0
0
3
0
Appendix
Table 10: F2 route-wise Hybrid correct counts, recomputed from archived per-question records. Agreement is defined by answer groups rather than correctness.
Parseable candidates
Agreement problems
Correct under Hybrid
1
14
6
2
7
5
3
11
10
4
54
53
Appendix
Table 11: F2 agreement states by the number of parseable initial candidates. SC@4 and ReSolve share the skipped outputs; all 86 added Uniform checks leave correctness unchanged. Counts are descriptive.
Method
AIME24
AIME25
BRUMO25
CMIMC25
HMMT25
Total
Oracle@4
27/26
25/25
24/24
30/30
20/20
126/125
SC@4
23/23
24/24
24/24
29/29
19/19
119/119
ReSolve
26/26
26/26
24/24
30/30
21/21
127/127
Matched operation
25/25
26/26
24/24
31/31
20/20
126/126
Uniform, closed loop
26/26
25/25
24/24
31/31
20/20
126/126
Operation permutation 1
24/24
26/26
24/24
32/32
20/20
126/126
Appendix
Table 12: H1 training-free results as Hybrid/CV counts ( N=160 ). Task sizes are 30, 30, 30, 40, and 30. All rows share the same four-candidate pool.
Configuration
H/CV
Calls
Input/q (K)
Output/q (K)
Total/q (K)
SC@4
119/119
0
0.00
0.00
87.67
ReSolve
127/127
74
2.01
4.56
94.24
Uniform, closed loop
126/126
68
1.78
4.36
93.80
One specialized call, original gate
125/125
58
1.50
3.56
92.73
One Uniform call, original gate
126/126
58
1.50
3.74
92.90
One Uniform call, validity gate
126/126
86
1.71
5.82
95.20
Appendix
Table 13: D1 on the H1 pool ( N=160 ). Input/output columns count moderation only; total includes proposals. Calls are panel totals and tokens are per-question means. The original gate triggers on 58 questions.
Large language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such sampling explores only through local decoding noise, it tends to produce many near duplicate attempts rather than genuinely different ideas. We ask whether exploration can instead be steered at a semantic level, by first sampling problem specific concepts, hints, or strategies and then conditioning answer generation on them. We refine this into a simple, more exploratory procedure that emits many diverse concepts in a single trajectory, and evaluate it on hard problems where repeated sampling struggles. We then go a step further and make concept generation trainable: a small concept generator is optimized with reinforcement learning so that its concepts maximize the downstream success of a larger, frozen answer generator. On hard mathematical reasoning problems, the trained concept generator substantially improves the answer generator's pass@k over naive repeated sampling at the same answer generation allocation, surpasses concepts drawn from much larger untuned models, and transfers to answer generators it was never trained against, including a model from a different family. A small model can thus be trained into an effective, reusable search policy for a much larger one.
Ismail Labiad, Matthieu Kowalski, Marc Schoenauer +2
Meta FAIR · Université Paris-Saclay, LISN, Inria, CNRS · NYU Courant Institute and CDS
Large Reasoning Models (LRMs) achieve strong performance on mathematical reasoning tasks but remain unreliable on challenging instances. Existing test-time scaling methods, such as repeated sampling, self-correction, and tree search, improve performance at the cost of increased computation, yet often exhibit diminishing returns on hard problems. We observe that output disagreement is strongly correlated with instance difficulty and prediction correctness, providing a useful signal for guiding instance-level strategy selection at test time. Based on this insight, we propose a training-free framework that formulates test-time scaling as an instance-level routing problem, rather than allocating more computation within a single strategy, dynamically selecting among different scaling strategies based on output disagreement. The framework applies lightweight resolution for consistent cases, majority voting for moderate disagreement, and rewriting-based reformulation for highly ambiguous instances. Experiments on seven mathematical benchmarks and three models show that our method improves accuracy by 3% - 7% while reducing sampling cost compared to existing approaches.
Zhimin Lin, Yixin Ji, Jinpeng Li +5
School of Computer Science and Technology, Soochow University · Department of Foundation Model, 2012 Labs, Huawei · 3Harbin Institute of Technology, Shenzhen (HITSZ)
Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing answer patterns instead of adding useful reasoning diversity. Verifier-based selection offers an alternative, but its performance depends on the calibration of an external reward model. We propose a verifier-free breadth--depth refinement framework that uses test-time compute to both explore and improve candidate solutions. The method samples multiple independent reasoning rollouts, refines each rollout through iterative self-critique and self-correction, and aggregates the refined answers by majority voting. Breadth preserves diverse initial attempts, while depth repairs local reasoning errors before aggregation. Across AIME24, AIME25, AMC, OlympiadBench, and MATH500, our method consistently improves over greedy decoding, majority voting, verifier-based best-of-N, beam search, and lookahead decoding across multiple open-weight models. For instance, with Qwen2.5-1.5B, accuracy increases from the strongest verifier-based baseline to 58.0% on MATH500, and from 25.0% to 32.5% on AMC. These results show that test-time compute can be more effective when used to refine sampled trajectories rather than only to sample more candidates or rely on verifier-guided selection.
Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer +4
University of Oklahoma · Stanford University · Universitat Pompeu Fabra +1