Recent thinking models are capable of solving complex reasoning tasks by scaling test-time compute, but this scaling should be allocated in line with task difficulty. On one hand, short reasoning (underthinking) leads to errors on harder problems that require extended reasoning steps; but, excessively long reasoning (overthinking) can be token-inefficient by generating unnecessary steps even after reaching a correct intermediate solution. We refer to this as under-adaptivity, where the model fails to modulate its response length appropriately given problems of varying difficulty. To address under-adaptivity and strike a balance between under- and overthinking, we propose TRAAC (Think Right with Adaptive, Attentive Compression), an online post-training RL method that leverages the model's self-attention to identify key steps and prune redundant ones. TRAAC also estimates difficulty and incorporates it into training rewards, thereby learning to allocate a reasoning budget commensurate with example difficulty. Across a variety of tasks (AIME, AMC, GPQA-D, BBEH), TRAAC (Qwen3-4B) achieves an average absolute accuracy gain of 8.4% with a relative reduction in reasoning length of 36.8% compared to the base model, and a 7.9% accuracy gain paired with a 29.4% length drop compared to the best RL baseline. TRAAC generalizes well, with accuracy and efficiency gains on out-of-distribution non-math datasets like GPQA-D, BBEH, and OptimalThinkingBench. Our analysis shows that TRAAC learns to adjust its thinking budget based on difficulty and that a combination of task-difficulty calibration and attention-based compression yields gains across diverse tasks.
Figures & tables
Figure 1: Overthinking on easy problems wastes tokens despite being able to maintain decent accuracy. On the other hand, underthinking on hard problems saves token budgets but fails to maintain accuracy. TRAAC addresses this trade-off by adapting to problem difficulty (estimated during training), via attention-based compression and, enabling intelligent resource allocation while improving both accuracy and efficiency.
Figure 2: Overview of TRAAC . Given a problem, the model first generates N rollouts, and the pass rate estimates the problem’s difficulty (easy, medium, or hard). Next, the generated reasoning is fed back into the model, which is asked to compute the attention score of each reasoning token from </think> ; we then remove steps with lower scores. The degree of removal is determined by the estimated difficulty: easier problems undergo more aggressive compression. Finally, we compute the correctness and length rewards over a group sampled from both the original and compressed rollouts, and use these rewards to update the policy.
Method
AIME
AMC
GPQA-D
BBEH
Average
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Qwen3-4B
Base Model
27.64
9.2
68.19
7.0
45.18
7.6
18.28
6.7
39.8
7.6
TokenSkip
5.84
9.6
27.71
8.7
32.32
7.8
11.91
7.2
19.4
8.3
L1-Max
30.11
7.1
63.61
5.8
43.23
5.8
14.91
5.0
38.0
5.9
LC-R1
13.48
2.6
56.38
1.7
26.67
1.5
12.35
1.9
27.2
1.9
Table 1: Performance comparison of TRAAC with various baselines. Each model is evaulate across various benchmarks, and Acc: accuracy(%) and Len: average Response Length(k) are reported. TRAAC on average shows the highest performance gain.
Method
OverthinkingBench
UnderthinkingBench
OptimalThinkingBench
Acc. ↑
Len. ↓
AUC↑
Acc. ↑
Len. ↓
F1 ↑
Qwen3-4B
Base Model
90.02
1.2
80.06
34.33
7.1
48.05
TokenSkip
78.15
3.5
57.88
14.80
7.9
23.57
L1-Max
87.22
0.9
1.11
21.27
6.3
2.10
LC-R1
78.62
0.3
64.20
14.95
1.3
24.25
Table 2: Performance of TRAAC and various baselines on OptimalThinkingBench. For UnderthinkingBench we report the Acc: Accuracy(%), and Len: Average Response length(k). For OverthinkingBench, in addition to Acc. and Len. we also report the AUCOAA .
Figure 3: (a) Relative change in compression rate of TRAAC and Qwen3-4B + Compression compared to Qwen3-4B across varying problem difficulty. (b) Absolute accuracy drop of TRAAC and Qwen3-4B + Compression compared to Qwen3-4B across varying difficulty.
Method
AIME
AMC
GPQA-D
BBEH
Avg.
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Qwen3-4B
Base Model
27.64
9.2
68.19
7.0
45.18
7.6
18.28
6.7
39.8
7.6
+ CR
44.36
7.9
77.35
5.5
46.29
5.7
18.13
5.2
46.5
6.1
+ LR
37.84
4.5
77.35
2.4
44.06
2.3
18.57
2.1
44.5
2.8
+ Compression
38.37
8.1
75.90
5.5
46.40
6.2
18.41
5.4
44.8
6.3
Table 3: Ablation Results of TRAAC (Qwen3-4B, Deepseek-Qwen-7B) tested across 4 datasets: AIME, AMC, GPQA-D, and BBEH. Each component addition adds to the previous method.
Method
OverthinkingBench
UnderthinkingBench
OptimalThinkingBench
Acc. ↑
Len. ↓
AUC↑
Acc. ↑
Len. ↓
F1 ↑
Qwen3-4B
Base Model
90.02
1.2
80.06
34.33
7.1
48.1
+ CR
90.02
0.9
78.86
37.06
5.7
50.4
+ LR
90.94
0.4
75.86
29.62
2.3
42.6
+ Compression
90.12
0.9
80.41
36.51
6.0
50.2
Table 4: Ablation Results of TRAAC (Qwen3-4B and Deepseek-Qwen-7B) on OptimalThinkingBench. Each component addition adds to the previous method.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Rollout (sec)
Optimise Policy (sec)
Total Time (sec)
Hardware
Base Model + CR
250
87.5
397.5
H100
Base Model + CR + LR
222
88
375
H100
TRAAC
418
88
583
H100
Appendix
Table 5: Training time breakdown for TRAAC and RL baselines during the first GRPO training step.
Method
FLOPs Used
Base Model + CR
1.65×1015 FLOPs
TRAAC
3.84×1015 FLOPs
Appendix
Table 6: FLOPs comparison for generating 20 training examples using different rollout strategies.
Figure 4: Time per step across training (Deepseek-Qwen-7B)
Method
Total FLOPs (80 questions)
Average FLOPs per question
Base Model + CR (Qwen3-4B)
3.7×1015 FLOPs
4.6×1013 FLOPs
TRAAC (Qwen3-4B)
2.7×1015 FLOPs
3.3×1013 FLOPs
Appendix
Table 7: Inference compute comparison for TRAAC vs. Base Model + CR on 80 AMC questions.
AIME
AMC
GPQA-D
Qwen3-4B
47.74 / 12.3
77.11 / 8.5
49.64 / 8.6
TRAAC
51.93 / 9.7
81.68 / 6.6
51.27 / 6.2
Appendix
Table 8: TRAAC with 15k training and test-time response length. For each dataset, Accuracy (%) and Response Length (in × 1000 tokens) are reported.
Pruning Strategy
AIME
AMC
GPQA-D
Random Steps
29.54 / 6.5
66.74 / 4.1
42.94 / 3.2
Least Confidence
32.35 / 5.8
71.08 / 3.4
47 / 3.0
TRAAC
45.45 / 6.7
79.52 / 4.2
47.2 / 4.2
Appendix
Table 9: Ablation on Qwen3-4B: comparing TRAAC with pruning random and least confident steps. For each dataset, Accuracy(%) / Response length (k) is reported.
Method
Avg. coherence score
Randomly pruned trajectories vs. Original Trajectories
3.35
TRAAC’s attention-based trajectories vs. Original Trajectories
3.71
Appendix
Table 10: Coherence comparison of attention based pruning strategy vs randomly pruned trajectory.
Task
Method
k=1
k=2
k=3
k=4
k=5
SR
Len
SR
Len
SR
Len
SR
Len
SR
Len
code_generation
Base Model
0.74
0.5
58.09
3.7
58.09
4.8
59.56
5.7
59.56
7.1
TRAAC
49.26
1.7
58.82
2.6
56.62
3.3
59.56
3.5
58.82
3.8
decision_making
Base Model
0.00
0.5
11.19
2.4
17.16
2.5
30.60
2.5
33.58
3.0
TRAAC
0.00
0.5
8.21
1.0
21.64
1.2
35.07
1.3
40.30
1.6
reasoning
Base Model
19.94
0.5
76.58
1.9
80.38
2.2
79.75
2.5
79.75
2.4
Appendix
Table 11: MINT benchmark results for Base Model (Qwen3-4B) and TRAAC across interaction limits k∈{1,2,3,4,5} . Metrics include success rate (SR, %) and average response length.
Method
AIME
AMC
GPQA
Acc.
Len.
Acc.
Len.
Acc.
Len.
TokenSkip
5.84%
9.6k
27.71%
8.7k
32.32%
7.8k
Base Model
27.64%
9.2k
68.19%
7.0k
45.18%
7.6k
SFT
26.06%
8.8k
59.51%
6.6k
42.00%
6.9k
TRAAC
45.45%
6.7k
79.52%
4.2k
47.21%
4.2k
Appendix
Table 12: Comparison of TokenSkip, Base Model, SFT on attention-compressed rollouts, and TRAAC across AIME, AMC, and GPQA. Metrics include accuracy (%) and average response length.
Method
AIME
AMC
GPQA
Acc.
Len.
Acc.
Len.
Acc.
Len.
TRAAC (full attention rollout)
2.309%
1.9k
18.53%
2.1k
25.81%
4.2k
TRAAC
45.45%
6.7k
79.52%
4.2k
47.21%
4.2k
Appendix
Table 13: Comparison of TRAAC with full-rollout attention pruning vs. standard TRAAC.
Method
AIME
AMC
GPQA-D
Average
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Phi-4-mini-reasoning
Base Model
31.4
8.5
67.4
5.9
40.3
7.7
46.4
7.4
+ CR
33.7
8.0
66.3
5.7
38.0
7.6
46.0
7.1
+ LR
32.5
8.0
66.3
5.5
42.1
7.2
47.0
6.9
TRAAC
28.3
7.7
70.1
5.1
44.5
5.6
47.6
6.1
Appendix
Table 14: Results of TRAAC on Phi-4-mini-reasoning tested across three datasets: AIME, AMC and GPQA-D. Each component addition adds to the previous method.
AIME
AMC
GPQA-D
TRAACreduced correctness
29.96 / 6.2
71.32 / 3.9
47.7 / 3.6
TRAACnon adaptive
5.33 / 0.6
34.87 / 0.7
29.79 / 0.5
TRAAC
45.45 / 6.7
79.52 / 4.2
47.21 / 4.2
Appendix
Table 15: Reward ablation results comparing different correctness and length reward configurations. For each dataset, Accuracy (%) and Response Length (in × 1000 tokens) are reported.
Method
AIME
AMC
GPQA-D
Base
27.64 / 9.2
68.19 / 7.0
45.18 / 7.6
Naive Linear Reward
25.30 / 6.6
64.09 / 4.2
39.80 / 3.2
Flattened Sigmoid ( c=0.5 )
31.91 / 6.4
71.32 / 4.0
44.87 / 2.8
TRAAC ( c=0.1 )
45.45 / 6.7
79.52 / 4.2
47.21 / 4.2
Appendix
Table 16: Ablation on the length reward smoothing using Qwen3-4B. For each dataset Accuracy(%) / Response length (k) are reported.
Method
AIME
AMC
GPQA-D
Base
27.64 / 9.2
68.19 / 7.0
45.18 / 7.6
Test-Time
32.13 / 8.5
70.60 / 5.7
47.61 / 6.6
TRAAC
45.45 / 6.7
79.52 / 4.2
47.20 / 4.2
Appendix
Table 17: Results of TRAAC as a test time method using Qwen3-4B, compared to base model and TRAAC . For each dataset Accuracy(%) / Response length (k) are reported.
Method
AIME
AMC
GPQA-D
BBEH
Average
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
Acc. ↑
Len. ↓
DeepSeek-R1-Distill-Qwen-7B
Base Model
33.71
8.2
74.22
5.7
43.55
7.1
10.61
5.9
40.5
6.7
+ CR
35.81
7.6
78.55
4.9
45.99
6.1
11.74
5.1
43.0
5.9
+ LR
32.73
6.0
79.04
3.3
45.99
3.5
11.51
2.7
42.3
3.9
TRAAC
38.60
7.3
77.83
4.5
47.31
6.2
11.55
5.2
43.8
5.8
Appendix
Table 18: Ablation Results of TRAAC on Deepseek-Qwen-7B tested across 4 datasets: AIME, AMC, GPQA-D, and BBEH. Each component addition adds to the previous method.
Method
OverthinkingBench
UnderthinkingBench
OptimalThinkingBench
Acc. ↑
Len. ↓
AUCOAA↑
Acc. ↑
Len. ↓
F1 ↑
DeepSeek-R1-Distill-Qwen-7B
Base Model
78.45
0.9
72.38
12.69
6.2
21.6
+ CR
79.51
0.8
73.36
17.05
5.7
27.7
+ LR
78.06
0.4
72.61
14.69
3.0
24.4
TRAAC
81.81
1.0
72.89
22.30
5.9
34.1
Appendix
Table 19: Ablation Results of TRAAC (Deepseek-Qwen-7B) on OptimalThinkingBench. Each component addition adds to the previous method.
Category
Hyperparameter
Value
Training
Number of rollouts
8
Temperature
1.0
top_p
1.0
top_k
-1.0
Max response length
10k
clip_ratio_low
0.20
Appendix
Table 20: Hyperparameters used for training, evaluation, and difficulty calibration.
Large reasoning models (LRMs) achieve strong performance via extended chain-of-thought (CoT) reasoning, yet suffer from excessive token consumption and high inference latency. Existing reinforcement learning (RL) approaches for CoT compression rely on uniform, static length penalties that neglect model capability dynamics and problem-level difficulty variation. We propose \textbf{ExpThink}\xspace, an RL framework that addresses both dimensions through two complementary mechanisms. First, \emph{experience-guided reward shaping} tracks the shortest correct solution found so far for each problem and applies a three-tier reward: full credit for concise correct responses, discounted credit for verbose correct ones, and zero for incorrect ones. The threshold tightens automatically with model improvement, forming a self-evolving curriculum that requires no manual scheduling. Second, \emph{difficulty-adaptive advantage} replaces standard deviation normalization with correct-count normalization, yielding monotonically difficulty-scaled gradients that amplify learning on hard problems to preserve accuracy while suppressing gradients on easy ones to encourage brevity. Together, these mechanisms enforce an accuracy-first, compression-second training objective. Experiments on multiple mathematical reasoning benchmarks demonstrate that \textbf{ExpThink}\xspace reduces average response length by up to 77% while simultaneously improving accuracy, achieving up to 3× higher accuracy-efficiency ratio (accuracy divided by average token count) than the vanilla baseline and outperforming existing RL-based compression methods on both metrics.
Tingcheng Bian, Yuzhe Zhang, Jing Jin +5
1Baidu Inc. · 2Shenzhen University · 3Peking University +2
Large Reasoning Models (LRMs) often overthink easy problems and underthink hard ones, leading to inefficient computation allocation. Existing methods regulate generated computation or select between direct answering and explicit reasoning, but do not jointly control whether}to reason and how much computation to allocate within reasoning. We call the resulting difficulty-dependent loss in accuracy under computation reduction the efficiency tax. We propose When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation. Its core mechanism, Instance-level Difficulty-Aware Control (IDAC), uses cached reference statistics of success and token cost to modulate a correctness-gated efficiency bonus based on generated token count. Importance sampling supports exploration of Think and NoThink, while Batch-Wise Standardization constructs standardized advantages for critic-free optimization. The framework requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates. On AIME24, When2Think improves Pass@3 by 10.0 percentage points while reducing token usage by 27.9% relative to the backbone.
Large reasoning models, such as OpenAI o1 and DeepSeek-R1, tend to become increasingly verbose as their reasoning capabilities improve. These inflated Chain-of-Thought (CoT) trajectories often exceed what the underlying problems require, wasting compute, latency, and context budgets. While introducing length-based efficiency rewards during reinforcement learning offers a natural remedy, existing methods struggle with two fundamental challenges: the optimal balance between correctness and efficiency is non-stationary throughout training, and intrinsic reasoning budgets vary drastically across problems. Relying on static reward weights and global length constraints inevitably forces a compromise between degraded accuracy and unrealized compression. To overcome these limitations, we propose LEAD (Length-Efficient Adaptive and Dynamic reasoning), a method that replaces static heuristics with online, self-adaptive mechanisms. LEAD dynamically calibrates the correctness-efficiency trade-off at each step using a Potential-Scaled Instability, directing optimization capacity to the most informative learning signal. Furthermore, it estimates an adaptive per-problem target length online based on the model's own correct rollouts, applying a symmetric efficiency reward that penalizes both overthinking and over-compression. Evaluated on five mathematical reasoning benchmarks, LEAD achieves the highest accuracy and Accuracy-Efficiency Score among RL-trained efficient-reasoning methods while producing substantially shorter outputs than the base model.
Songtao Wei, Yi Li, Zhikai Li +7
University of Texas at Dallas · Emory University · University of Texas at Arlington +2