Recent thinking models are capable of solving complex reasoning tasks by scaling test-time compute, but this scaling should be allocated in line with task difficulty. On one hand, short reasoning (underthinking) leads to errors on harder problems that require extended reasoning steps; but, excessively long reasoning (overthinking) can be token-inefficient by generating unnecessary steps even after reaching a correct intermediate solution. We refer to this as under-adaptivity, where the model fails to modulate its response length appropriately given problems of varying difficulty. To address under-adaptivity and strike a balance between under- and overthinking, we propose TRAAC (Think Right with Adaptive, Attentive Compression), an online post-training RL method that leverages the model's self-attention to identify key steps and prune redundant ones. TRAAC also estimates difficulty and incorporates it into training rewards, thereby learning to allocate a reasoning budget commensurate with example difficulty. Across a variety of tasks (AIME, AMC, GPQA-D, BBEH), TRAAC (Qwen3-4B) achieves an average absolute accuracy gain of 8.4% with a relative reduction in reasoning length of 36.8% compared to the base model, and a 7.9% accuracy gain paired with a 29.4% length drop compared to the best RL baseline. TRAAC generalizes well, with accuracy and efficiency gains on out-of-distribution non-math datasets like GPQA-D, BBEH, and OptimalThinkingBench. Our analysis shows that TRAAC learns to adjust its thinking budget based on difficulty and that a combination of task-difficulty calibration and attention-based compression yields gains across diverse tasks.
Figures & tables
Figure 1: Overthinking on easy problems wastes tokens despite being able to maintain decent accuracy. On the other hand, underthinking on hard problems saves token budgets but fails to maintain accuracy. TRAAC addresses this trade-off by adapting to problem difficulty (estimated during training), via attention-based compression and, enabling intelligent resource allocation while improving both accuracy and efficiency.
Figure 2: Overview of TRAAC . Given a problem, the model first generates N rollouts, and the pass rate estimates the problem’s difficulty (easy, medium, or hard). Next, the generated reasoning is fed back into the model, which is asked to compute the attention score of each reasoning token from </think> ; we then remove steps with lower scores. The degree of removal is determined by the estimated difficulty: easier problems undergo more aggressive compression. Finally, we compute the correctness and length rewards over a group sampled from both the original and compressed rollouts, and use these rewards to update the policy.
Method
AIME
AMC
GPQA-D
BBEH
Average
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Qwen3-4B
Base Model
27.64
9.2
68.19
7.0
45.18
7.6
18.28
6.7
39.8
7.6
TokenSkip
5.84
9.6
27.71
8.7
32.32
7.8
11.91
7.2
19.4
8.3
L1-Max
30.11
7.1
63.61
5.8
43.23
5.8
14.91
5.0
38.0
5.9
LC-R1
13.48
2.6
56.38
1.7
26.67
1.5
12.35
1.9
27.2
1.9
Table 1: Performance comparison of TRAAC with various baselines. Each model is evaulate across various benchmarks, and Acc: accuracy(%) and Len: average Response Length(k) are reported. TRAAC on average shows the highest performance gain.
Method
OverthinkingBench
UnderthinkingBench
OptimalThinkingBench
Acc. ↑
Len. ↓
AUC↑
Acc. ↑
Len. ↓
F1 ↑
Qwen3-4B
Base Model
90.02
1.2
80.06
34.33
7.1
48.05
TokenSkip
78.15
3.5
57.88
14.80
7.9
23.57
L1-Max
87.22
0.9
1.11
21.27
6.3
2.10
LC-R1
78.62
0.3
64.20
14.95
1.3
24.25
Table 2: Performance of TRAAC and various baselines on OptimalThinkingBench. For UnderthinkingBench we report the Acc: Accuracy(%), and Len: Average Response length(k). For OverthinkingBench, in addition to Acc. and Len. we also report the AUCOAA .
Figure 3: (a) Relative change in compression rate of TRAAC and Qwen3-4B + Compression compared to Qwen3-4B across varying problem difficulty. (b) Absolute accuracy drop of TRAAC and Qwen3-4B + Compression compared to Qwen3-4B across varying difficulty.
Method
AIME
AMC
GPQA-D
BBEH
Avg.
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Qwen3-4B
Base Model
27.64
9.2
68.19
7.0
45.18
7.6
18.28
6.7
39.8
7.6
+ CR
44.36
7.9
77.35
5.5
46.29
5.7
18.13
5.2
46.5
6.1
+ LR
37.84
4.5
77.35
2.4
44.06
2.3
18.57
2.1
44.5
2.8
+ Compression
38.37
8.1
75.90
5.5
46.40
6.2
18.41
5.4
44.8
6.3
Table 3: Ablation Results of TRAAC (Qwen3-4B, Deepseek-Qwen-7B) tested across 4 datasets: AIME, AMC, GPQA-D, and BBEH. Each component addition adds to the previous method.
Method
OverthinkingBench
UnderthinkingBench
OptimalThinkingBench
Acc. ↑
Len. ↓
AUC↑
Acc. ↑
Len. ↓
F1 ↑
Qwen3-4B
Base Model
90.02
1.2
80.06
34.33
7.1
48.1
+ CR
90.02
0.9
78.86
37.06
5.7
50.4
+ LR
90.94
0.4
75.86
29.62
2.3
42.6
+ Compression
90.12
0.9
80.41
36.51
6.0
50.2
Table 4: Ablation Results of TRAAC (Qwen3-4B and Deepseek-Qwen-7B) on OptimalThinkingBench. Each component addition adds to the previous method.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Rollout (sec)
Optimise Policy (sec)
Total Time (sec)
Hardware
Base Model + CR
250
87.5
397.5
H100
Base Model + CR + LR
222
88
375
H100
TRAAC
418
88
583
H100
Appendix
Table 5: Training time breakdown for TRAAC and RL baselines during the first GRPO training step.
Method
FLOPs Used
Base Model + CR
1.65×1015 FLOPs
TRAAC
3.84×1015 FLOPs
Appendix
Table 6: FLOPs comparison for generating 20 training examples using different rollout strategies.
Figure 4: Time per step across training (Deepseek-Qwen-7B)
Method
Total FLOPs (80 questions)
Average FLOPs per question
Base Model + CR (Qwen3-4B)
3.7×1015 FLOPs
4.6×1013 FLOPs
TRAAC (Qwen3-4B)
2.7×1015 FLOPs
3.3×1013 FLOPs
Appendix
Table 7: Inference compute comparison for TRAAC vs. Base Model + CR on 80 AMC questions.
AIME
AMC
GPQA-D
Qwen3-4B
47.74 / 12.3
77.11 / 8.5
49.64 / 8.6
TRAAC
51.93 / 9.7
81.68 / 6.6
51.27 / 6.2
Appendix
Table 8: TRAAC with 15k training and test-time response length. For each dataset, Accuracy (%) and Response Length (in × 1000 tokens) are reported.
Pruning Strategy
AIME
AMC
GPQA-D
Random Steps
29.54 / 6.5
66.74 / 4.1
42.94 / 3.2
Least Confidence
32.35 / 5.8
71.08 / 3.4
47 / 3.0
TRAAC
45.45 / 6.7
79.52 / 4.2
47.2 / 4.2
Appendix
Table 9: Ablation on Qwen3-4B: comparing TRAAC with pruning random and least confident steps. For each dataset, Accuracy(%) / Response length (k) is reported.
Method
Avg. coherence score
Randomly pruned trajectories vs. Original Trajectories
3.35
TRAAC’s attention-based trajectories vs. Original Trajectories
3.71
Appendix
Table 10: Coherence comparison of attention based pruning strategy vs randomly pruned trajectory.
Task
Method
k=1
k=2
k=3
k=4
k=5
SR
Len
SR
Len
SR
Len
SR
Len
SR
Len
code_generation
Base Model
0.74
0.5
58.09
3.7
58.09
4.8
59.56
5.7
59.56
7.1
TRAAC
49.26
1.7
58.82
2.6
56.62
3.3
59.56
3.5
58.82
3.8
decision_making
Base Model
0.00
0.5
11.19
2.4
17.16
2.5
30.60
2.5
33.58
3.0
TRAAC
0.00
0.5
8.21
1.0
21.64
1.2
35.07
1.3
40.30
1.6
reasoning
Base Model
19.94
0.5
76.58
1.9
80.38
2.2
79.75
2.5
79.75
2.4
Appendix
Table 11: MINT benchmark results for Base Model (Qwen3-4B) and TRAAC across interaction limits k∈{1,2,3,4,5} . Metrics include success rate (SR, %) and average response length.
Method
AIME
AMC
GPQA
Acc.
Len.
Acc.
Len.
Acc.
Len.
TokenSkip
5.84%
9.6k
27.71%
8.7k
32.32%
7.8k
Base Model
27.64%
9.2k
68.19%
7.0k
45.18%
7.6k
SFT
26.06%
8.8k
59.51%
6.6k
42.00%
6.9k
TRAAC
45.45%
6.7k
79.52%
4.2k
47.21%
4.2k
Appendix
Table 12: Comparison of TokenSkip, Base Model, SFT on attention-compressed rollouts, and TRAAC across AIME, AMC, and GPQA. Metrics include accuracy (%) and average response length.
Method
AIME
AMC
GPQA
Acc.
Len.
Acc.
Len.
Acc.
Len.
TRAAC (full attention rollout)
2.309%
1.9k
18.53%
2.1k
25.81%
4.2k
TRAAC
45.45%
6.7k
79.52%
4.2k
47.21%
4.2k
Appendix
Table 13: Comparison of TRAAC with full-rollout attention pruning vs. standard TRAAC.
Method
AIME
AMC
GPQA-D
Average
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Phi-4-mini-reasoning
Base Model
31.4
8.5
67.4
5.9
40.3
7.7
46.4
7.4
+ CR
33.7
8.0
66.3
5.7
38.0
7.6
46.0
7.1
+ LR
32.5
8.0
66.3
5.5
42.1
7.2
47.0
6.9
TRAAC
28.3
7.7
70.1
5.1
44.5
5.6
47.6
6.1
Appendix
Table 14: Results of TRAAC on Phi-4-mini-reasoning tested across three datasets: AIME, AMC and GPQA-D. Each component addition adds to the previous method.
AIME
AMC
GPQA-D
TRAACreduced correctness
29.96 / 6.2
71.32 / 3.9
47.7 / 3.6
TRAACnon adaptive
5.33 / 0.6
34.87 / 0.7
29.79 / 0.5
TRAAC
45.45 / 6.7
79.52 / 4.2
47.21 / 4.2
Appendix
Table 15: Reward ablation results comparing different correctness and length reward configurations. For each dataset, Accuracy (%) and Response Length (in × 1000 tokens) are reported.
Method
AIME
AMC
GPQA-D
Base
27.64 / 9.2
68.19 / 7.0
45.18 / 7.6
Naive Linear Reward
25.30 / 6.6
64.09 / 4.2
39.80 / 3.2
Flattened Sigmoid ( c=0.5 )
31.91 / 6.4
71.32 / 4.0
44.87 / 2.8
TRAAC ( c=0.1 )
45.45 / 6.7
79.52 / 4.2
47.21 / 4.2
Appendix
Table 16: Ablation on the length reward smoothing using Qwen3-4B. For each dataset Accuracy(%) / Response length (k) are reported.
Method
AIME
AMC
GPQA-D
Base
27.64 / 9.2
68.19 / 7.0
45.18 / 7.6
Test-Time
32.13 / 8.5
70.60 / 5.7
47.61 / 6.6
TRAAC
45.45 / 6.7
79.52 / 4.2
47.20 / 4.2
Appendix
Table 17: Results of TRAAC as a test time method using Qwen3-4B, compared to base model and TRAAC . For each dataset Accuracy(%) / Response length (k) are reported.
Method
AIME
AMC
GPQA-D
BBEH
Average
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
DeepSeek-R1-Distill-Qwen-7B
Base Model
33.71
8.2
74.22
5.7
43.55
7.1
10.61
5.9
40.5
6.7
+ CR
35.81
7.6
78.55
4.9
45.99
6.1
11.74
5.1
43.0
5.9
+ LR
32.73
6.0
79.04
3.3
45.99
3.5
11.51
2.7
42.3
3.9
TRAAC
38.60
7.3
77.83
4.5
47.31
6.2
11.55
5.2
43.8
5.8
Appendix
Table 18: Ablation Results of TRAAC on Deepseek-Qwen-7B tested across 4 datasets: AIME, AMC, GPQA-D, and BBEH. Each component addition adds to the previous method.
Method
OverthinkingBench
UnderthinkingBench
OptimalThinkingBench
Acc. ↑
Len. ↓
AUCOAA↑
Acc. ↑
Len. ↓
F1 ↑
DeepSeek-R1-Distill-Qwen-7B
Base Model
78.45
0.9
72.38
12.69
6.2
21.6
+ CR
79.51
0.8
73.36
17.05
5.7
27.7
+ LR
78.06
0.4
72.61
14.69
3.0
24.4
TRAAC
81.81
1.0
72.89
22.30
5.9
34.1
Appendix
Table 19: Ablation Results of TRAAC (Deepseek-Qwen-7B) on OptimalThinkingBench. Each component addition adds to the previous method.
Category
Hyperparameter
Value
Training
Number of rollouts
8
Temperature
1.0
top_p
1.0
top_k
-1.0
Max response length
10k
clip_ratio_low
0.20
Appendix
Table 20: Hyperparameters used for training, evaluation, and difficulty calibration.