Large language models (LLMs) exhibiting test-time scaling behavior, such as extended reasoning traces and self-verification, have demonstrated remarkable performance on complex, long-term reasoning tasks. However, the robustness of these reasoning behaviors remains underexplored. To investigate this, we conduct a systematic evaluation of multiple reasoning models across three scenarios: (1) problems augmented with lengthy, irrelevant context; (2) multi-turn conversational settings with independent tasks; and (3) problems presented as a subtask within a complex task. We observe an interesting phenomenon: reasoning models tend to produce much shorter reasoning traces (up to 74%) for the same problem under different context conditions compared to the traces produced when the problem is presented in isolation. A finer-grained analysis reveals that this compression is associated with a decrease in self-verification and uncertainty management behaviors, such as double-checking. Importantly, we show that even when additional self-checks are forced, their efficiency depends not only on the content of the reasoning traces, but also on the presence of redundant context. We hope our findings draw additional attention to both the robustness of reasoning models and the problem of context management for LLMs.
Figures & tables
Model
Baseline
Subtask
Long Input
Multi-turn
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
IMOAnswerBench
Qwen3.5-27B
74.5
28,771
62.4
20,165
67.8
16,415
67.0
17,404
Gemma 4 31B
72.7
9,240
52.8
5,635
67.0
6,252
71.0
7,530
gpt-oss-120b
73.8
24,180
64.0
17,408
64.0
11,876
69.3
19,831
Gemini 3 Flash Preview
82.8
23,090
67.0
13,653
80.3
19,879
82.5
21,693
Table 1: Model performance on IMOAnswerBench and GPQA-Diamond. Accuracy and average number of generated reasoning tokens are shown. Background color represents the relative change from the baseline values.
Figure 1: Accuracy comparison between Baseline and Long Input conditions across task difficulty levels (Easy, Mid, Hard) on a subset of IMOAnswerBench. Left: Gemma 4 31B, right: gpt-oss-120b.
Figure 2: Average reasoning length on MATH-500 under varying number of inserted tokens in Long Input and Multi-turn setups. Left: Gemma 4 31B, right: Qwen3.5-27B.
Figure 3: Number of generated tokens for Qwen3.5-27B for each MATH-500 task. X-axis: Baseline, Y-Axis: Long Input.
Model
Baseline
Subtask
Long Input
Multi-turn
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Olmo-3-7B-Instruct
91.2
925
89.1
1,139
89.8
1,162
90.5
1,089
Olmo-3-7B-Think-SFT
94.5
2,671
90.2
1,829
91.4
1,908
89.1
1,611
Olmo-3-7B-Think-DPO
93.5
3,051
89.7
2,286
91.0
2,437
88.7
2,395
Olmo-3-7B-Think
95.2
3,664
92.1
2,851
93.2
2,859
92.2
2,788
Table 2: Model performance with accuracy and average number of generated reasoning tokens, MATH-500. For Instruct model, the average number of response tokens is reported. Background color represents the relative change from the baseline values.
Token
0 ( Baseline )
128
16k
</think>
21%
26%
46%
Wait
11%
10%
5%
Alternatively
17%
11%
5%
But
46%
38%
20%
Maybe
23%
17%
9%
Table 3: Resampling from the same reasoning traces but under varying numbers of inserted prompt tokens in the Long Input setup. The table presents the ratio of traces containing the end of reasoning or self-verification tokens.
Trace
Baseline
Long Input(128)
Long Input(16k)
Multi-turn
Baseline
62.0%
68.0%
90.4%
78.4%
Long Input
41.2%
48.2%
85.4%
64.4%
Multi-turn
46.8%
55.5%
86.7%
70.5%
Table 4: The same reasoning traces receive higher self-confidence scores under non-baseline context conditions. Table represents the ratio of traces with the highest self-confidence scores. Rows represent how traces were generated, columns represent the context used during self-confidence evaluation. Qwen3-32B, MATH-500.
Long Input
“Wait” intervention (Long Input context)
“Wait” intervention (Baseline context)
Baseline
Model
Acc.
Tokens
Acc.
Tokens
Acc.
Tokens
Acc.
Tokens
Gemma 4 31B
67.0
6,252
68.3
8,058
71.8
10,301
72.7
9,240
Qwen3.5-27B
67.8
16,415
68.5
18,191
69.3
19,865
74.5
28,771
Table 5: Effect of a “Wait” intervention on IMOAnswerBench accuracy and average reasoning-token count. The same reasoning traces generated under Long Input are continued after the intervention with the irrelevant context either retained or removed.
Baseline
Long Input
Model
Prompt
Acc.
Tokens
Acc.
Tokens
Qwen3.5-27B
Standard
74.5
28,771
67.8
16,415
Max effort
74.4
28,755
70.5
15,960
Gemma 4 31B
Standard
72.7
9,240
67.0
6,252
Max effort
73.8
10,624
70.9
7,198
gpt-oss-120b
Standard
73.8
24,180
64.0
11,876
Table 6: Effect of max reasoning effort prompt on IMOAnswerBench.
Model
Baseline
Multi-turn(16)
Multi-turn(32)
Long Input(32)
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Olmo-3-7B-Think-SFT
94.5
2,671
92.2
1,781
89.1
1,611
91.4
1,908
Olmo-3-7B-Think-SFT-Ours
96.1
2,869
94.5
2,674
89.8
2,606
93.8
2,806
Table 7: Model performance with accuracy and average number of generated reasoning tokens under different context conditions on MATH-500.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Input
Output
Qwen3.5-27B
$0.195 / M
$1.56 / M
Gemma 4 31B
$0.14 / M
$0.40 / M
gpt-oss-120b
$0.039 / M
$0.18 / M
Gemini 3 Flash Preview
$0.50 / M
$3.00 / M
Kimi K2 Thinking
$0.60 / M
$2.50 / M
Appendix
Table 8: API pricing used for estimating evaluation costs. Prices are reported in USD per million tokens, following the OpenRouter model pages at the time of the experiments.
Baseline
Long Input (Shakespeare)
Long Input (IMO lemmas)
Model
Acc.
Tokens
Acc.
Tokens
Acc.
Tokens
gpt-oss-120b
73.8
24,180
64.0
11,876
63.5
11,055
Qwen3.5-27B
74.5
28,771
67.8
16,415
68.0
15,465
Appendix
Table 9: Model performance on IMOAnswerBench under Long Input scenario with different content used as prefix.
Model
Baseline
Subtask
Long Input
Multi-turn
Acc.
Tok.
Acc1
Acc2
Tok.
Acc.
Tok.
Acc.
Tok.
Qwen3.5-27B
87
22,837
78
78
16,302
85
17,429
81
6,222
gpt-oss-120b
89
10,666
51
86
7,675
88
7,173
83
6,156
Appendix
Table 10: Model performance with accuracy and average number of generated reasoning tokens on LiveCodeBench. For the Subtask setup, accuracies are reported separately for the two generated subtasks. Background color represents the relative change from the baseline values.
Model
Baseline
First subproblem
Second subproblem
Qwen3.5-27B
74.5
66.8
58.0
gpt-oss-120b
73.8
63.8
64.3
Gemini 3 Flash Preview
82.8
68.3
65.8
Kimi K2 Thinking
74.8
68.0
62.0
Appendix
Table 11: Model accuracy on each subproblem in the Subproblem setup. Background color represents the relative change from the baseline values.
Trace
Baseline
Long Input(128)
Long Input(16k)
Multi-turn
Baseline
85.9%
88.4%
90.4%
88.8%
Long Input
82.1%
87.1%
90.2%
86.7%
Multi-turn
83.0%
86.5%
89.0%
86.7%
Appendix
Table 12: Verbalized confidence experiment with last 64 tokens removed, using numeric scores. The same reasoning traces receive higher self-confidence scores under non-baseline context conditions. Table represents the ratio of traces with the highest self-confidence scores. Rows represent how traces were generated, columns represent the context used during self-confidence evaluation. Qwen3-32B, MATH-500.
Trace
Baseline
Long Input(128)
Long Input(16k)
Multi-turn
Baseline
38.0%
50.3%
57.8%
48.6%
Long Input
16.8%
28.3%
36.8%
26.8%
Multi-turn
22.5%
33.7%
44.8%
31.8%
Appendix
Table 13: Verbalized confidence experiment with the first half of the trace being evaluated. The same reasoning traces receive higher self-confidence scores under non-baseline context conditions. Table represents the ratio of traces with the highest self-confidence scores. Rows represent how traces were generated, columns represent the context used during self-confidence evaluation. Qwen3-32B, MATH-500.
Model
Baseline
Long Input
Multi-turn
Qwen3-32B
Total reasoning tokens
3,421
2,375
2,825
First candidate answer position
797
698
753
Qwen3.5-27B
Total reasoning tokens
4,689
2,545
3,017
First candidate answer position
574
516
533
Gemma-4-31B
Total reasoning tokens
1,929
1,227
1,391
First candidate answer position
656
638
649
Appendix
Table 14: Average reasoning length and first candidate answer position across models and settings, MATH-500.
Model
Baseline
Long Input
Multi-turn
Multi-turn(4)
Qwen3.5-27B
Baseline
12.2
14.8
18.8
28.8
Long Input
3.1
11.0
11.3
14.1
Multi-turn
5.3
5.7
11.1
20.1
Gemma-4-31B
Baseline
35.0
35.5
56.8
60.7
Long Input
16.1
16.1
32.9
34.2
Multi-turn
16.9
20.7
33.5
39.3
Appendix
Table 15: The ratio of finished samples during resampling experiment. Rows represent the source of the trace, column - the context condition during resampling. The same almost finished reasoning traces tend to finish earlier under non-baseline conditions. Multi-turn(4) represents the increased chat history - with 4 interactions.
Figure 4: Difference of transition probability matrices (Long Input - Baseline). Qwen3-32B, MATH-500 problems.
Model
GPQA-Diamond
MMLU-Pro
Acc.
Tok.
Acc.
Tok.
Olmo-3-7B-Think-SFT
42.6
9,222
49.2
1779
Olmo-3-7B-Think-SFT-Ours
43.1
9,164
54.2
1912
Appendix
Table 16: GPQA-Diamond and MMLU-Pro performance of the original Olmo-3-7B-Think-SFT model and the fine-tuned variant. Tokens denote the average number of generated reasoning tokens.
Figure 5: Number of generated tokens for each IMOAnswerBench task. X-axis: Baseline, Y-Axis: Long Input.