Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to solve in a single attempt. In contrast, sequential methods build on previous answers to access new ideas, yet so far have not been shown to reach answers beyond those found by parallel scaling. First, we show that sequential scaling often stops improving because it becomes prematurely trapped in an attractor: a set of answers that prevents exploration of different answers once entered. Across 27 combinations of scaling methods, models, and benchmarks, we find that 53.8% of sequential scaling trajectories enter an attractor within four iterations. Second, we show that a simple model-mixing intervention helps escape attractors. This reduces the attractor hit rate by 21.2 percentage points on average, expands solution coverage beyond a compute-matched parallel baseline, and improves accuracy of recursive self-aggregation by at least 2.2 percentage points. Our results motivate refocusing long-horizon test-time scaling from parallel methods to sequential methods that improve previous answers.
Figures & tables
Figure 1: Escaping locally optimal attractors allows test-time scaling to improve for longer. Left : Schematic illustration of two sequential scaling trajectories through the state space, with minimum points as attractors. Both trajectories initially reach an attractor; the baseline (blue throughout this paper) remains trapped, while our intervention (orange throughout this paper) escapes and continues exploring. Right : Coverage (problems solved at any point during scaling) on AMO-Bench-P, over time. Our method continues improving after the baseline saturates by mixing Qwen3 235B A22B Thinking 2507 with GPT OSS 120B.
Table 1: How different test-time scaling methods fit into our general formulation.
Attractor hit by S4 (%) ↓
Different correctness (%) ↓
Method
Model
IMO
AMO
RGG
IMO
AMO
RGG
Reasoning Cache
GPT OSS 120B
50.1
39.3
23.3
15.1
21.0
9.5
Gemma 3 12B IT
48.0
37.6
23.7
17.1
19.4
20.4
Granite 4.1 8B
73.2
62.4
35.3
22.2
20.1
27.4
Self-Refine
GPT OSS 120B
57.1
37.6
27.0
15.5
14.3
10.5
Gemma 3 12B IT
48.0
33.3
33.3
16.7
17.4
10.3
Table 2: Left columns : Probability that each scaling method hits an attractor within the first 4 states. Higher values mean the scaling-method–model combination is more likely to get stuck early on. Right columns : Probability that, if two trajectories – using the same scaling method, model, and question – both terminate in attractors, then exactly one of the attractors contains only incorrect answers, making it a suboptimal attractor.
Figure 2: Illustration of how mixing models can help escape suboptimal attractors. Y axis represents attractor strength (i.e. attractors that satisfy Definition 4.1 for lower values of ϵ0 and ϵ1 are stronger).
Attractor hit by S4 (%) ↓
Different correctness (%) ↓
Method
Model
IMO
AMO
RGG
IMO
AMO
RGG
Reasoning Cache
GPT
50.1
39.3
23.3
15.1
21.0
9.5
Mmix (ours)
30.3
11.1
12.7
15.1
14.3
10.5
Self-Refine
GPT
57.1
37.6
27.0
15.5
14.3
10.5
Mmix (ours)
33.4
11.1
18.7
8.8
20.4
6.7
Recursive Self-Aggregation
GPT
88.7
72.6
81.0
12.7
9.8
12.7
Table 3: Comparison of attractor behaviour between scaling methods using Mmix and methods using only GPT OSS 120B, for the same experiments as in Table 2 . Left columns : Probability that the scaling method hits an attractor within the first 4 states for a single trajectory. Right columns : Probability that, if two trajectories – using the same scaling method, model, and question – both terminate in attractors, then exactly one of the attractors contains only incorrect answers, making it a suboptimal attractor.
Figure 3: Coverage over time of different sequential methods using Mmix across three benchmarks. Mmix (orange) outperforms both the sequential baseline that uses GPT OSS 120B (blue), and the parallel baseline (grey) in almost every case.
Figure 4: Coverage for alternative model mixtures for RSA on IMO-AnswerBench. Left: Results for different combinations of models, to identify whether some models are better at escaping attractors than others. Right: Results for different numbers of non-primary models where the primary model is always selected with probability \nicefrac13 , to test whether adding more models provides meaningfully more opportunities for one model being able to escape attractors.
Figure 5: Effect of mixture concentration c on RSA with Mmix over 10 states. Left : Coverage is maximised by occasionally sampling weaker models. Centre : Probability of generating a correct answer when there was one in the previous state improves as sampling concentrates on the primary model. Right : Using a single model outperforms mixtures of models overall.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Coverage before and after entering an attractor on IMO-AnswerBench, with trajectories aligned to the state where they enter the attractor (grey dashed line). Progress slows substantially after entry, indicating that we identify attractor states that limit further exploration.
Method
Model
Attractor hit by S4 (%) ↓
Reasoning Cache
GPT OSS 120B
47.67±2.33
Gemma 3 12B IT
46.67±1.33
Granite 4.1 8B
73.50±0.50
Self-Refine
GPT OSS 120B
57.31±0.33
Gemma 3 12B IT
44.33±3.67
Granite 4.1 8B
69.17±2.17
Appendix
Table 4: Sensitivity of attractor hitting-time probabilities to randomness in Algorithm 1 for IMO-AnswerBench (i.e. standard deviation for left column of Table 2 ). Results are stable across two different seeds used in Algorithm 1 .
Figure 7: Probability of entering an attractor at or before state index x for IMO-AnswerBench. The probability begins to level off after four states, motivating the four-state threshold used throughout our experiments.
Different attractors (%) ↓
Method
Model
IMO
AMO
RGG
Reasoning Cache
GPT OSS 120B
33.7
22.2
31.0
Gemma 3 12B IT
34.3
18.9
32.2
Granite 4.1 8B
68.0
65.3
63.7
Mmix (ours)
20.5
20.0
8.2
Self-Refine
GPT OSS 120B
47.3
47.6
23.7
Appendix
Table 5: Probability that, if we generate two trajectories using the same scaling method, model, and question, and both terminate in attractors, then the two attractors are different. Higher values suggest there are more attractors in the state space.
Figure 8: Example of Mmix escaping an attractor containing only incorrect answers by using a non-primary model to help revise its reasoning and reach the correct answer (69169).
Figure 9: Per-problem coverage changes between Mmix and GPT OSS 120B with different sequential scaling methods. Mmix improves over GPT OSS 120B without losing many correct answers.
Figure 10: Coverage achieved by each scaling-method–model combination against price on OpenRouter (2026) (identical to Figure 3 , but with cost on the x axis).
Coverage (%) ↑
Model
IMO
AMO
RGG
GPT OSS 120B
82.8
64.5
79.0
Mmix (ours)
82.0
62.8
77.3
Appendix
Table 6: Parallel scaling coverage over 60 answers. The Mmix row is identical to results displayed in Figure 3 . Sampling from Mmix alone does not improve coverage over GPT OSS 120B, so its sequential scaling gains in Figure 3 are not explained by complementary model capabilities.
Figure 11: Coverage achieved by recursive self-aggregation with GPT OSS 120B (medium reasoning effort) on IMO-AnswerBench across different temperature parameters, using an identical configuration to Figure 3 .
Figure 12: Results of mixing Qwen 3.5 122B A10B with GPT OSS 120B (both medium reasoning effort) in recursive self-aggregation, for IMO-AnswerBench. Except for the models used, this experiment is identical to those in Figure 4 .
Figure 13: Comparison between recursive self-aggregation with a single model, and with annealed model sampling. For annealed model sampling, we initialise c as 0.75 to select the model for the first answer, and then increase c linearly at each attempt, until we have c=1 for the final attempt. We use an otherwise identical experimental setup to that in Figure 5 .
Method
Model
Attractor hit by S4 (%) ↓
Different attractors (%) ↓
Reasoning Cache
GPT OSS 120B
20.6
23.6
Gemma 3 12B IT
13.3
16.4
Granite 4.1 8B
25.9
53.4
Mmix (ours)
16.4
34.8
Self-Refine
GPT OSS 120B
8.0
25.0
Gemma 3 12B IT
7.7
30.0
Appendix
Table 7: Attractor dynamics on ARC AGI 1, similar to Tables 2 and 3 . Sequential scaling is less likely to get stuck in attractors within the first few states compared to the other benchmarks tested.
Figure 14: Coverage over time on ARC AGI 1 (similar to Figure 3 ). The GPT OSS 120B baseline continues improving after the first few states, outperforming Mmix .
Figure 15: Per-problem agreement between Mmix and GPT OSS 120B on ARC AGI 1 (similar to Figure 9 ).
Test-time scaling improves language model reasoning by spending additional compute to explore multiple solution trajectories. The key challenge is to maximize accuracy while minimizing the total number of generated tokens during reasoning. Recent PRM-guided methods score intermediate prefixes to steer this search, but most are frontier-only: they keep only the current active prefixes and irreversibly prune or resample away the rest using noisy PRM scores. This can cause premature commitment, diversity collapse, and the loss of prefixes that still admit correct continuations. We introduce stochastic backtracking over a persistent pool of historical prefixes, allowing test-time compute to revisit previously generated states instead of only expanding the current frontier. To make this efficient, we propose two complementary mechanisms. Subpool Selection strengthens greedy PRM-guided search by applying Top-N selection within random subpools, giving historical prefixes a chance to bypass over-scored frontier candidates. Power Backtrack Sequential Monte Carlo extends SMC-style resampling to the persistent pool using powered PRM scores and mixture-corrected weights. Across mathematical reasoning benchmarks and model scales, our methods consistently achieve higher accuracy per token count, and the same level of accuracy using only a fraction of the token count in comparison to strong PRM-guided baselines, demonstrating that persistent-pool stochastic backtracking provides a simple and effective way to improve the accuracy-token trade-off in test-time scaling.
Two forms of test-time scaling for Large Language Models (LLMs) have emerged as effective and widely adopted paradigms: sequential, in which later answer attempts depend on earlier ones, and parallel, such as i.i.d. sampling with reranking. In this study, we investigate their properties in translation. First, our study shows that sequential sampling has a higher performance ceiling, providing a more diverse and effective pool of samples, particularly under smaller sampling budgets. Second, we interrogate the nature of test-time scaling through a multidimensional manual analysis. Human analysis of the Best-of-N translations demonstrates that sequential sampling substantially improves translation fluency and naturalness, but can degrade accuracy when inference budgets are large. Finally, we suggest an explanation of the mechanism through which sequential scaling improves machine translation. Our controlled analysis partially attributes the success of sequential self-improvement to the model's access to a larger target-side context. Ablation experiments on sequential sampling demonstrate its robustness across different sampling temperatures, while also revealing sensitivity to context construction, suggesting directions for future improvement.
Di Wu, Sergey Troshin, Christof Monz +2
University of Amsterdam · Vrije Universiteit Amsterdam
Test-time scaling has become an effective paradigm for improving the reasoning ability of large language models by allocating additional computation during inference. Recent structured approaches have further advanced this paradigm by organizing inference across multiple trajectories, refinement rounds, and verification-based feedback. However, existing structured test-time scaling methods either weakly coordinate parallel reasoning trajectories or rely on noisy historical information without explicitly deciding what should be retained and reused, limiting their ability to balance exploration and exploitation. In this work, we propose TMAS, a framework for scaling test-time compute via multi-agent synergy. TMAS organizes inference as a collaborative process among specialized agents, enabling structured information flow across agents, trajectories, and refinement iterations. To support effective cross-trajectory collaboration, TMAS introduces hierarchical memories: the experience bank reuses low-level reliable intermediate conclusions and local feedback, while the guideline bank records previously explored high-level strategies to steer subsequent rollouts away from redundant reasoning patterns. Furthermore, we design a hybrid reward reinforcement learning scheme tailored to TMAS, which jointly preserves basic reasoning capability, enhances experience utilization, and encourages exploration beyond previously attempted solution strategies. Extensive experiments on challenging reasoning benchmarks show that TMAS achieves stronger iterative scaling than existing test-time scaling baselines, with hybrid reward training further improving scaling effectiveness and stability across iterations. Code and data are available at https://github.com/IQuestLab/tmas.