Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to solve in a single attempt. In contrast, sequential methods build on previous answers to access new ideas, yet so far have not been shown to reach answers beyond those found by parallel scaling. First, we show that sequential scaling often stops improving because it becomes prematurely trapped in an attractor: a set of answers that prevents exploration of different answers once entered. Across 27 combinations of scaling methods, models, and benchmarks, we find that 53.8% of sequential scaling trajectories enter an attractor within four iterations. Second, we show that a simple model-mixing intervention helps escape attractors. This reduces the attractor hit rate by 21.2 percentage points on average, expands solution coverage beyond a compute-matched parallel baseline, and improves accuracy of recursive self-aggregation by at least 2.2 percentage points. Our results motivate refocusing long-horizon test-time scaling from parallel methods to sequential methods that improve previous answers.
Figures & tables
Figure 1: Escaping locally optimal attractors allows test-time scaling to improve for longer. Left : Schematic illustration of two sequential scaling trajectories through the state space, with minimum points as attractors. Both trajectories initially reach an attractor; the baseline (blue throughout this paper) remains trapped, while our intervention (orange throughout this paper) escapes and continues exploring. Right : Coverage (problems solved at any point during scaling) on AMO-Bench-P, over time. Our method continues improving after the baseline saturates by mixing Qwen3 235B A22B Thinking 2507 with GPT OSS 120B.
Table 1: How different test-time scaling methods fit into our general formulation.
Attractor hit by S4 (%) ↓
Different correctness (%) ↓
Method
Model
IMO
AMO
RGG
IMO
AMO
RGG
Reasoning Cache
GPT OSS 120B
50.1
39.3
23.3
15.1
21.0
9.5
Gemma 3 12B IT
48.0
37.6
23.7
17.1
19.4
20.4
Granite 4.1 8B
73.2
62.4
35.3
22.2
20.1
27.4
Self-Refine
GPT OSS 120B
57.1
37.6
27.0
15.5
14.3
10.5
Gemma 3 12B IT
48.0
33.3
33.3
16.7
17.4
10.3
Table 2: Left columns : Probability that each scaling method hits an attractor within the first 4 states. Higher values mean the scaling-method–model combination is more likely to get stuck early on. Right columns : Probability that, if two trajectories – using the same scaling method, model, and question – both terminate in attractors, then exactly one of the attractors contains only incorrect answers, making it a suboptimal attractor.
Figure 2: Illustration of how mixing models can help escape suboptimal attractors. Y axis represents attractor strength (i.e. attractors that satisfy Definition 4.1 for lower values of ϵ0 and ϵ1 are stronger).
Attractor hit by S4 (%) ↓
Different correctness (%) ↓
Method
Model
IMO
AMO
RGG
IMO
AMO
RGG
Reasoning Cache
GPT
50.1
39.3
23.3
15.1
21.0
9.5
Mmix (ours)
30.3
11.1
12.7
15.1
14.3
10.5
Self-Refine
GPT
57.1
37.6
27.0
15.5
14.3
10.5
Mmix (ours)
33.4
11.1
18.7
8.8
20.4
6.7
Recursive Self-Aggregation
GPT
88.7
72.6
81.0
12.7
9.8
12.7
Table 3: Comparison of attractor behaviour between scaling methods using Mmix and methods using only GPT OSS 120B, for the same experiments as in Table 2 . Left columns : Probability that the scaling method hits an attractor within the first 4 states for a single trajectory. Right columns : Probability that, if two trajectories – using the same scaling method, model, and question – both terminate in attractors, then exactly one of the attractors contains only incorrect answers, making it a suboptimal attractor.
Figure 3: Coverage over time of different sequential methods using Mmix across three benchmarks. Mmix (orange) outperforms both the sequential baseline that uses GPT OSS 120B (blue), and the parallel baseline (grey) in almost every case.
Figure 4: Coverage for alternative model mixtures for RSA on IMO-AnswerBench. Left: Results for different combinations of models, to identify whether some models are better at escaping attractors than others. Right: Results for different numbers of non-primary models where the primary model is always selected with probability \nicefrac13 , to test whether adding more models provides meaningfully more opportunities for one model being able to escape attractors.
Figure 5: Effect of mixture concentration c on RSA with Mmix over 10 states. Left : Coverage is maximised by occasionally sampling weaker models. Centre : Probability of generating a correct answer when there was one in the previous state improves as sampling concentrates on the primary model. Right : Using a single model outperforms mixtures of models overall.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Coverage before and after entering an attractor on IMO-AnswerBench, with trajectories aligned to the state where they enter the attractor (grey dashed line). Progress slows substantially after entry, indicating that we identify attractor states that limit further exploration.
Method
Model
Attractor hit by S4 (%) ↓
Reasoning Cache
GPT OSS 120B
47.67±2.33
Gemma 3 12B IT
46.67±1.33
Granite 4.1 8B
73.50±0.50
Self-Refine
GPT OSS 120B
57.31±0.33
Gemma 3 12B IT
44.33±3.67
Granite 4.1 8B
69.17±2.17
Appendix
Table 4: Sensitivity of attractor hitting-time probabilities to randomness in Algorithm 1 for IMO-AnswerBench (i.e. standard deviation for left column of Table 2 ). Results are stable across two different seeds used in Algorithm 1 .
Figure 7: Probability of entering an attractor at or before state index x for IMO-AnswerBench. The probability begins to level off after four states, motivating the four-state threshold used throughout our experiments.
Different attractors (%) ↓
Method
Model
IMO
AMO
RGG
Reasoning Cache
GPT OSS 120B
33.7
22.2
31.0
Gemma 3 12B IT
34.3
18.9
32.2
Granite 4.1 8B
68.0
65.3
63.7
Mmix (ours)
20.5
20.0
8.2
Self-Refine
GPT OSS 120B
47.3
47.6
23.7
Appendix
Table 5: Probability that, if we generate two trajectories using the same scaling method, model, and question, and both terminate in attractors, then the two attractors are different. Higher values suggest there are more attractors in the state space.
Figure 8: Example of Mmix escaping an attractor containing only incorrect answers by using a non-primary model to help revise its reasoning and reach the correct answer (69169).
Figure 9: Per-problem coverage changes between Mmix and GPT OSS 120B with different sequential scaling methods. Mmix improves over GPT OSS 120B without losing many correct answers.
Figure 10: Coverage achieved by each scaling-method–model combination against price on OpenRouter (2026) (identical to Figure 3 , but with cost on the x axis).
Coverage (%) ↑
Model
IMO
AMO
RGG
GPT OSS 120B
82.8
64.5
79.0
Mmix (ours)
82.0
62.8
77.3
Appendix
Table 6: Parallel scaling coverage over 60 answers. The Mmix row is identical to results displayed in Figure 3 . Sampling from Mmix alone does not improve coverage over GPT OSS 120B, so its sequential scaling gains in Figure 3 are not explained by complementary model capabilities.
Figure 11: Coverage achieved by recursive self-aggregation with GPT OSS 120B (medium reasoning effort) on IMO-AnswerBench across different temperature parameters, using an identical configuration to Figure 3 .
Figure 12: Results of mixing Qwen 3.5 122B A10B with GPT OSS 120B (both medium reasoning effort) in recursive self-aggregation, for IMO-AnswerBench. Except for the models used, this experiment is identical to those in Figure 4 .
Figure 13: Comparison between recursive self-aggregation with a single model, and with annealed model sampling. For annealed model sampling, we initialise c as 0.75 to select the model for the first answer, and then increase c linearly at each attempt, until we have c=1 for the final attempt. We use an otherwise identical experimental setup to that in Figure 5 .
Method
Model
Attractor hit by S4 (%) ↓
Different attractors (%) ↓
Reasoning Cache
GPT OSS 120B
20.6
23.6
Gemma 3 12B IT
13.3
16.4
Granite 4.1 8B
25.9
53.4
Mmix (ours)
16.4
34.8
Self-Refine
GPT OSS 120B
8.0
25.0
Gemma 3 12B IT
7.7
30.0
Appendix
Table 7: Attractor dynamics on ARC AGI 1, similar to Tables 2 and 3 . Sequential scaling is less likely to get stuck in attractors within the first few states compared to the other benchmarks tested.
Figure 14: Coverage over time on ARC AGI 1 (similar to Figure 3 ). The GPT OSS 120B baseline continues improving after the first few states, outperforming Mmix .
Figure 15: Per-problem agreement between Mmix and GPT OSS 120B on ARC AGI 1 (similar to Figure 9 ).