Large language models (LLMs) exhibiting test-time scaling behavior, such as extended reasoning traces and self-verification, have demonstrated remarkable performance on complex, long-term reasoning tasks. However, the robustness of these reasoning behaviors remains underexplored. To investigate this, we conduct a systematic evaluation of multiple reasoning models across three scenarios: (1) problems augmented with lengthy, irrelevant context; (2) multi-turn conversational settings with independent tasks; and (3) problems presented as a subtask within a complex task. We observe an interesting phenomenon: reasoning models tend to produce much shorter reasoning traces (up to 74%) for the same problem under different context conditions compared to the traces produced when the problem is presented in isolation. A finer-grained analysis reveals that this compression is associated with a decrease in self-verification and uncertainty management behaviors, such as double-checking. Importantly, we show that even when additional self-checks are forced, their efficiency depends not only on the content of the reasoning traces, but also on the presence of redundant context. We hope our findings draw additional attention to both the robustness of reasoning models and the problem of context management for LLMs.
Figures & tables
Model
Baseline
Subtask
Long Input
Multi-turn
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
IMOAnswerBench
Qwen3.5-27B
74.5
28,771
62.4
20,165
67.8
16,415
67.0
17,404
Gemma 4 31B
72.7
9,240
52.8
5,635
67.0
6,252
71.0
7,530
gpt-oss-120b
73.8
24,180
64.0
17,408
64.0
11,876
69.3
19,831
Gemini 3 Flash Preview
82.8
23,090
67.0
13,653
80.3
19,879
82.5
21,693
Table 1: Model performance on IMOAnswerBench and GPQA-Diamond. Accuracy and average number of generated reasoning tokens are shown. Background color represents the relative change from the baseline values.
Figure 1: Accuracy comparison between Baseline and Long Input conditions across task difficulty levels (Easy, Mid, Hard) on a subset of IMOAnswerBench. Left: Gemma 4 31B, right: gpt-oss-120b.
Figure 2: Average reasoning length on MATH-500 under varying number of inserted tokens in Long Input and Multi-turn setups. Left: Gemma 4 31B, right: Qwen3.5-27B.
Figure 3: Number of generated tokens for Qwen3.5-27B for each MATH-500 task. X-axis: Baseline, Y-Axis: Long Input.
Model
Baseline
Subtask
Long Input
Multi-turn
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Olmo-3-7B-Instruct
91.2
925
89.1
1,139
89.8
1,162
90.5
1,089
Olmo-3-7B-Think-SFT
94.5
2,671
90.2
1,829
91.4
1,908
89.1
1,611
Olmo-3-7B-Think-DPO
93.5
3,051
89.7
2,286
91.0
2,437
88.7
2,395
Olmo-3-7B-Think
95.2
3,664
92.1
2,851
93.2
2,859
92.2
2,788
Table 2: Model performance with accuracy and average number of generated reasoning tokens, MATH-500. For Instruct model, the average number of response tokens is reported. Background color represents the relative change from the baseline values.
Token
0 ( Baseline )
128
16k
</think>
21%
26%
46%
Wait
11%
10%
5%
Alternatively
17%
11%
5%
But
46%
38%
20%
Maybe
23%
17%
9%
Table 3: Resampling from the same reasoning traces but under varying numbers of inserted prompt tokens in the Long Input setup. The table presents the ratio of traces containing the end of reasoning or self-verification tokens.
Trace
Baseline
Long Input(128)
Long Input(16k)
Multi-turn
Baseline
62.0%
68.0%
90.4%
78.4%
Long Input
41.2%
48.2%
85.4%
64.4%
Multi-turn
46.8%
55.5%
86.7%
70.5%
Table 4: The same reasoning traces receive higher self-confidence scores under non-baseline context conditions. Table represents the ratio of traces with the highest self-confidence scores. Rows represent how traces were generated, columns represent the context used during self-confidence evaluation. Qwen3-32B, MATH-500.
Long Input
“Wait” intervention (Long Input context)
“Wait” intervention (Baseline context)
Baseline
Model
Acc.
Tokens
Acc.
Tokens
Acc.
Tokens
Acc.
Tokens
Gemma 4 31B
67.0
6,252
68.3
8,058
71.8
10,301
72.7
9,240
Qwen3.5-27B
67.8
16,415
68.5
18,191
69.3
19,865
74.5
28,771
Table 5: Effect of a “Wait” intervention on IMOAnswerBench accuracy and average reasoning-token count. The same reasoning traces generated under Long Input are continued after the intervention with the irrelevant context either retained or removed.
Baseline
Long Input
Model
Prompt
Acc.
Tokens
Acc.
Tokens
Qwen3.5-27B
Standard
74.5
28,771
67.8
16,415
Max effort
74.4
28,755
70.5
15,960
Gemma 4 31B
Standard
72.7
9,240
67.0
6,252
Max effort
73.8
10,624
70.9
7,198
gpt-oss-120b
Standard
73.8
24,180
64.0
11,876
Table 6: Effect of max reasoning effort prompt on IMOAnswerBench.
Model
Baseline
Multi-turn(16)
Multi-turn(32)
Long Input(32)
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Olmo-3-7B-Think-SFT
94.5
2,671
92.2
1,781
89.1
1,611
91.4
1,908
Olmo-3-7B-Think-SFT-Ours
96.1
2,869
94.5
2,674
89.8
2,606
93.8
2,806
Table 7: Model performance with accuracy and average number of generated reasoning tokens under different context conditions on MATH-500.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Input
Output
Qwen3.5-27B
$0.195 / M
$1.56 / M
Gemma 4 31B
$0.14 / M
$0.40 / M
gpt-oss-120b
$0.039 / M
$0.18 / M
Gemini 3 Flash Preview
$0.50 / M
$3.00 / M
Kimi K2 Thinking
$0.60 / M
$2.50 / M
Appendix
Table 8: API pricing used for estimating evaluation costs. Prices are reported in USD per million tokens, following the OpenRouter model pages at the time of the experiments.
Baseline
Long Input (Shakespeare)
Long Input (IMO lemmas)
Model
Acc.
Tokens
Acc.
Tokens
Acc.
Tokens
gpt-oss-120b
73.8
24,180
64.0
11,876
63.5
11,055
Qwen3.5-27B
74.5
28,771
67.8
16,415
68.0
15,465
Appendix
Table 9: Model performance on IMOAnswerBench under Long Input scenario with different content used as prefix.
Model
Baseline
Subtask
Long Input
Multi-turn
Acc.
Tok.
Acc1
Acc2
Tok.
Acc.
Tok.
Acc.
Tok.
Qwen3.5-27B
87
22,837
78
78
16,302
85
17,429
81
6,222
gpt-oss-120b
89
10,666
51
86
7,675
88
7,173
83
6,156
Appendix
Table 10: Model performance with accuracy and average number of generated reasoning tokens on LiveCodeBench. For the Subtask setup, accuracies are reported separately for the two generated subtasks. Background color represents the relative change from the baseline values.
Model
Baseline
First subproblem
Second subproblem
Qwen3.5-27B
74.5
66.8
58.0
gpt-oss-120b
73.8
63.8
64.3
Gemini 3 Flash Preview
82.8
68.3
65.8
Kimi K2 Thinking
74.8
68.0
62.0
Appendix
Table 11: Model accuracy on each subproblem in the Subproblem setup. Background color represents the relative change from the baseline values.
Trace
Baseline
Long Input(128)
Long Input(16k)
Multi-turn
Baseline
85.9%
88.4%
90.4%
88.8%
Long Input
82.1%
87.1%
90.2%
86.7%
Multi-turn
83.0%
86.5%
89.0%
86.7%
Appendix
Table 12: Verbalized confidence experiment with last 64 tokens removed, using numeric scores. The same reasoning traces receive higher self-confidence scores under non-baseline context conditions. Table represents the ratio of traces with the highest self-confidence scores. Rows represent how traces were generated, columns represent the context used during self-confidence evaluation. Qwen3-32B, MATH-500.
Trace
Baseline
Long Input(128)
Long Input(16k)
Multi-turn
Baseline
38.0%
50.3%
57.8%
48.6%
Long Input
16.8%
28.3%
36.8%
26.8%
Multi-turn
22.5%
33.7%
44.8%
31.8%
Appendix
Table 13: Verbalized confidence experiment with the first half of the trace being evaluated. The same reasoning traces receive higher self-confidence scores under non-baseline context conditions. Table represents the ratio of traces with the highest self-confidence scores. Rows represent how traces were generated, columns represent the context used during self-confidence evaluation. Qwen3-32B, MATH-500.
Model
Baseline
Long Input
Multi-turn
Qwen3-32B
Total reasoning tokens
3,421
2,375
2,825
First candidate answer position
797
698
753
Qwen3.5-27B
Total reasoning tokens
4,689
2,545
3,017
First candidate answer position
574
516
533
Gemma-4-31B
Total reasoning tokens
1,929
1,227
1,391
First candidate answer position
656
638
649
Appendix
Table 14: Average reasoning length and first candidate answer position across models and settings, MATH-500.
Model
Baseline
Long Input
Multi-turn
Multi-turn(4)
Qwen3.5-27B
Baseline
12.2
14.8
18.8
28.8
Long Input
3.1
11.0
11.3
14.1
Multi-turn
5.3
5.7
11.1
20.1
Gemma-4-31B
Baseline
35.0
35.5
56.8
60.7
Long Input
16.1
16.1
32.9
34.2
Multi-turn
16.9
20.7
33.5
39.3
Appendix
Table 15: The ratio of finished samples during resampling experiment. Rows represent the source of the trace, column - the context condition during resampling. The same almost finished reasoning traces tend to finish earlier under non-baseline conditions. Multi-turn(4) represents the increased chat history - with 4 interactions.
Figure 4: Difference of transition probability matrices (Long Input - Baseline). Qwen3-32B, MATH-500 problems.
Model
GPQA-Diamond
MMLU-Pro
Acc.
Tok.
Acc.
Tok.
Olmo-3-7B-Think-SFT
42.6
9,222
49.2
1779
Olmo-3-7B-Think-SFT-Ours
43.1
9,164
54.2
1912
Appendix
Table 16: GPQA-Diamond and MMLU-Pro performance of the original Olmo-3-7B-Think-SFT model and the fine-tuned variant. Tokens denote the average number of generated reasoning tokens.
Figure 5: Number of generated tokens for each IMOAnswerBench task. X-axis: Baseline, Y-Axis: Long Input.
While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural entanglement in monolithic architectures, where the massive search burden of contextual grounding exhausts the representational capacity needed for complex reasoning. To resolve this, we propose Grounding-Reasoning Disaggregation via DIStributed long COntext scaling (DISCO). Inspired by distributed computing frameworks like Apache Spark, DISCO partitions long context across a fleet of Worker LLMs dedicated exclusively to parallel, localized grounding. A central Driver LLM, trained via Reinforcement Learning (GRPO) to optimize planning, orchestrates execution by dynamically mapping queries into atomic extraction tasks and reducing the gathered evidence to synthesize a final answer. By isolating reasoning from raw context noise, DISCO effectively eliminates context rot. On RULER-QA (1M tokens), it maintains 78.4% accuracy where standard baselines collapse. Furthermore, it outperforms full-context models by up to 9.8 points on LongBench v2 and matches frontier models like Gemini-3-Pro-Preview while reducing inference costs by over 80%, establishing a highly efficient paradigm for robust long-context inference.
Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient. In this work, we introduce Structured Thoughts, a framework that organizes reasoning into alternating <try> and <outcome> blocks: <try> captures exploratory scratch work, while <outcome> contains the distilled conclusion of that step. We construct a dataset of structured thoughts by segmenting reasoning traces into <try> blocks and prompting an LLM to summarize each step into its corresponding <outcome>. Fine-tuning pretrained foundation models on this reformatted data produces models that adopt the structured reasoning style, leading to performance gains of up to 8.08% on reasoning benchmarks compared to standard SFT. The explicit structure also enables context pruning: after each <try>/<outcome> pair, the <try> can be pruned, allowing the model to retain conclusions without keeping the full scratch work in the context. A proof-of-concept pruning implementation achieves an average of 85% memory / context savings with an 8.67% performance drop across mathematical tasks.
Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test-time compute. Improving reasoning quality directly would require process reward models, but the step-level annotations needed to train them are expensive and scarce. We find such a signal in how the model's confidence evolves during reasoning: premature confidence, the tendency to commit to an answer early and use the remaining tokens to rationalize it, strongly predicts flawed reasoning across tasks and model scales. We exploit this in progressive confidence shaping, a reinforcement learning objective that trains models to update their confidence as they reason rather than commit early -- rewarding gradual confidence growth and penalizing early commitment, with no external labels or reward models. The method improves accuracy and reasoning quality from 1.5B to 8B parameters across arithmetic (Countdown), math (DAPO, AIME), and science (ScienceQA): on Countdown, accuracy improves 3.2x (+42.0pp) and flawed reasoning drops 48pp; on AIME, Pass@64 improves 6.6pp. Consistent with this mechanism, the method also improves faithfulness: on a safety benchmark, our models more transparently surface misleading content in their reasoning traces rather than concealing it. Controlled experiments reveal that the problem and its remedy scale together: premature confidence grows with model size and task difficulty, and so do the gains from addressing it.
Jingchu Gai, Guanning Zeng, Christina Baek +4
1Carnegie Mellon University · 2Tsinghua University