Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this issue but inevitably truncate necessary steps and induce underthinking, thereby compromising performance. To address this dilemma, we investigate the latent representations and observe that efficient reasoning steps naturally cluster into a concentrated region in latent space, while those deviating from this region tend to produce verbose sequences. To leverage this, we keep reasoning focused within this region via a quadratic program which projects deviating hidden states back into the region. Then we propose a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance. Extensive experiments conducted on four models ranging from 1.5B to 14B, and across six benchmarks in math reasoning, coding, and scientific QA, validate the effectiveness of our method, up to a 12.1% improvement in accuracy while reducing generated tokens by 11.8% to 52.8%. Codes are available at \href{https://github.com/hzn18/Opt4Reasoning}{https://github.com/hzn18/Opt4Reasoning}.
Figures & tables
Figure 1 : Visualization of different reasoning steps in the latent space. The horizontal and vertical axes correspond to the first and second PCA components, respectively.
Figure 2 : (a) The distribution of distances to the set H for trajectories with correct and incorrect answers. (b) The relation between the distance d(c,H) and the total token length of the reasoning trajectory.
Methods
MATH
PROGRAMMING
SCIENCE
GSM8K
MATH500
AMC2023
AIME2025
LiveCodeBench
GPQA-Diamond
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
DeepSeek-R1-Distill-Qwen-1.5B
Vanilla
81.5
1489
82.1
4476
64.4
7708
17.3
11811
32.1
10013
33.3
8241
CCoT
76.6
577
82.0
3832
68.8
7012
22.7
11077
30.6
9824
33.3
7592
DEER
74.2
672
72.2
2578
62.5
5139
20.2
9601
23.7
8227
29.7
7655
Table 1 : Performance comparison against baselines. For each column within a model group, the best result is bolded and the second-best is underlined .
Figure 3 : The impact of the key hyperparameters for performance in MATH500: (a) Intervention Strength. (b) Steering Layer. (c) PCA dimension.
Method
# Tokens
TPS ↑
TPR (s) ↓
DeepSeek-R1-Distill-Qwen-1.5B
Baseline
4678
5596
0.83
Ours (Online)
3074
4682
0.66 (-21.2%)
Ours (Explicit)
2812
5472
0.51 (-38.4%)
DeepSeek-R1-Distill-Qwen-7B
Baseline
3631
2659
1.36
Table 2 : Efficiency Analysis on A800 GPU.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : Visualization of different reasoning steps generated by Qwen3-4B. The horizontal and vertical axes correspond to the first and second PCA components, respectively.
Figure 5 : Visualization of different reasoning steps generated by DeepSeek-R1-Distill-Qwen-7B. The horizontal and vertical axes correspond to the first and second PCA components, respectively.
Figure 6 : Visualization of different reasoning steps generated by Qwen3-14B. The horizontal and vertical axes correspond to the first and second PCA components, respectively.
Figure 7 : Visualization of latent representations across different PCA dimensions. The results show that Direct CoT steps remain concentrated in a core region across various dimension pairs (e.g., Dim 0 vs 1, Dim 2 vs 3), while overthinking behaviors such as Transition and Reflection steps fall into the scattered region.
Methods
GSM8K
MATH500
AMC2023
AIME2025
LiveCodeBench
GPQA-Diamond
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Acc.
Tok.
Vanilla
81.5
1489
82.1
4476
64.4
7708
17.3
11811
32.1
10013
33.3
8241
Ours-MATH500
83.1
853
83.2
2949
75.2
4821
26.7
8502
32.7
7656
36.4
5651
Ours-GSM8K
82.3
707
82.3
2980
76.7
4825
23.3
8097
32.8
7575
35.8
5738
Ours-AMC23
82.6
782
82.9
2951
72.5
4609
22.2
8649
32.5
7594
35.3
5590
Ours-AIME2025
82.7
901
82.0
3294
71.6
4939
30.0
8727
33.4
8393
34.9
6149
Appendix
Table 3 : Cross-Domain and Cross-Difficulty Transferability on DeepSeek-R1-Distill-Qwen-1.5B.
Method
Reflection
Transition
Performance
Word Count
TF
TF–IDF
Word Count
TF
TF–IDF
Acc.
Tok.
DeepSeek-R1-Distill-Qwen-1.5B
Baseline
28.5
9.3
13.0
8.0
2.5
4.7
82.1
4476
Ours
6.0
4.2
6.6
1.7
0.7
1.5
83.2
2949
DeepSeek-R1-Distill-Qwen-7B
Baseline
18.3
7.6
10.3
4.9
2.1
3.7
91.6
3647
Appendix
Table 4 : Semantic statistics of Transition and Reflection vocabularies.
Table 6 : Pass@k Performance on AMC 2023 and AIME 2025.
Model
Method
AMC2023
AIME2025
Acc
Tokens
Acc
Tokens
DeepSeek-1.5B
Vanilla
64.4 ± 5.4
7708 ± 703
17.3 ± 3.7
11811 ± 625
Ours
75.2 ± 3.5
4821 ± 304
26.7 ± 4.2
8502 ± 390
DeepSeek-7B
Vanilla
86.3 ± 4.0
5846 ± 358
26.9 ± 3.9
11209 ± 780
Ours
91.1 ± 2.7
4083 ± 261
39.0 ± 4.0
9042 ± 631
Qwen-4B
Vanilla
95.9 ± 1.7
7588 ± 475
60.2 ± 3.7
16835 ± 944
Appendix
Table 7: Statistical reliability evaluation on AMC23 and AIME2025. We report mean accuracy, mean token count and corresponding standard deviations across datasets and model sizes.
Model
Variables
Constraints
Online (ms)
Explicit (ms)
Additional Memory (MB)
R1-1.5B
10
8
1.7
0.3
0.41
R1-7B
12
8
2.3
0.3
0.93
Qwen3-4B
12
8
2.9
0.3
0.56
Qwen3-14B
16
8
5.3
0.5
1.80
Appendix
Table 8 : Comparison of Average Solving Time. We detail the scale of the QP problem and compare the latency of the online solver with the explicit acceleration. The additional memory denotes the GPU footprint required to store the precomputed piecewise linear mappings for the explicit method.
Symbol
Description
Xdir,Xvan
The original representations of Direct CoTs and Vanilla CoTs.
Zdir,Zvan
The representations of Direct CoTs and Vanilla CoTs after PCA projection.
H
The efficient reasoning.
U
The PCA projection matrix.
k
The dimension of the latent space after the PCA projection.
M
The number of clusters defining the efficient reasoning set H .
Appendix
Table 9 : Mathematical notations and their descriptions used in our method.
Figure 8 : The impact of intervention strength λ on Pass@1 accuracy and average token count across three models on the MATH500 dataset.
Figure 9 : The impact of the steering layer index on Pass@1 accuracy and average token count across three models on the MATH500 dataset.
Figure 10 : The impact of the PCA dimension k on Pass@1 accuracy and average token count across three models on the MATH500 dataset.
Figure 11 : The impact of the boundary shift proportion α on Pass@1 accuracy and average token count across three models on the MATH500 dataset.
Figure 12 : The impact of the cluster number M (number of linear constraints) on Pass@1 accuracy and average token count across three models on the MATH500 dataset