MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller
Authors: Zhibin Wen, Tao Han, Lei Bai, Can Li, Yang Xu
Organizations: Southern University of Science and Technology · The Hong Kong University of Science and Technology · Shanghai Artificial Intelligence Laboratory · Tongji University
Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Conversely, aggressively shortening reasoning can degrade performance on difficult ones. Effective reasoning therefore requires dynamically deciding when additional computation is useful based on the reasoner's capabilities and evolving solution state. Existing approaches often rely on predefined budgets or intervention rules, retrain the target reasoner, or require additional supervision. We introduce MetaCtrl, a lightweight controller that adaptively regulates a frozen reasoner without predefined token budgets or reasoner retraining. We formulate reasoning regulation as a sequential metacognitive control problem: MetaCtrl observes the evolving reasoning trace and decides whether to continue, simplify, skip redundant steps, or conclude reasoning. It is trained directly with reinforcement learning using a reward that prioritizes correctness while favoring shorter trajectories among correct solutions, requiring neither supervised intervention trajectories nor problem-specific budgets. Across seven benchmarks spanning mathematics, science, and code, MetaCtrl consistently improves the accuracy of LRMs while reducing their reasoning length. On DeepSeek-R1-Distill-Qwen-7B, it improves average accuracy by 4.7 points while reducing generation length by 53.3%. Without further training, the same controller transfers to an unseen reasoner (e.g., Qwen3-14B), improving average accuracy by 2.9 points and reducing generation length by 50.3%. These results establish MetaCtrl as a plug-and-play controller for improving reasoning accuracy while substantially reducing inference-time generation. The code is available at https://github.com/binbin2xs/MetaCtrl.
Figures & tables
Figure 1: Left: Accuracy and generation length of the same reasoner when controlled by MetaCtrl, controlled by frontier LLMs, or run without external control across five benchmarks. Right: Comparison of MetaCtrl with representative efficient reasoning baselines across two reasoning models.
Figure 2: Comparison of four paradigms for efficient LRMs inference. In external control methods, some approaches train the reasoner, while others keep it frozen.
Figure 3: Overview of MetaCtrl. (a) A learned meta-level controller regulates a frozen reasoner through four intervention actions, without an input budget. (b) At each turn, it selects an intervention based on the current reasoning state. (c) The controller is trained with GRPO using an efficiency-aware outcome reward. (d) The trained controller generalizes to unseen reasoners without retraining.
Math Reasoning
Scientific Reasoning
Methods
MATH-500
AIME2024
OmniMath
GPQA Diamond
AVG
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc. ↑
Reduc. ↑
DeepSeek-R1-Distill-Qwen-7B [Seen Reasoner]
NoThinking
79.4
699
-75.5
40.0
4134
-60.9
40.0
2083
-63.7
29.3
2084
-61.0
47.2
-65.3
Vanilla
85.2
2857
−0%
50.0
10570
−0%
45.0
5736
−0%
32.8
5349
−0%
53.3
−0%
D-Prompt
86.0
2726
−4.6%
46.7
10209
−3.4%
45.0
5398
−5.9%
30.3
5117
−4.3%
52.0
−4.6%
Table 1: Comparison across reasoning models of different scales and and multiple baselines.
Math Reasoning
Scientific Reasoning
Methods
MATH-500
AIME 2024
OlympiadBench
GPQA Diamond
AVG
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc. ↑
Reduc. ↑
Qwen3-8B [Unseen Reasoner]
NoThinking
87.4
1480
-70.0
23.3
7121
-41.2
48.7
5219
-43.7
48.5
2564
-72.7
52.0
-56.9
Vanilla
92.2
4926
−0%
63.3
12101
−0%
59.9
9268
−0%
52.5
9382
−0%
67.0
−0%
TALE
91.8
3682
−25.3%
60.0
11847
−2.1%
56.1
7306
−21.2%
52.0
4161
−55.6%
65.0
−26.0%
Table 2: Comparison across the Qwen3 series at different model scales and multiple baselines.
Controller Model
MATH-500
AMC
AIME2024
OlympiadBench
GPQA Diamond
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Vanilla (No Controller)
85.2
2857
73.5
7279
50.0
10570
48.6
8347
32.8
5349
Qwen-3.8-Flash
71.4
6548
+129.2%
60.2
9229
+26.8%
–
–
–
–
–
–
30.3
10584
+97.9%
Qwen-3.8-Max
78.2
1605
-43.8%
48.8
2507
-65.6%
30.0
7780
-26.4%
39.3
2921
-65.0%
38.4
1813
-66.1%
GLM-5.2
82.0
1472
-48.5%
66.9
2450
-66.3%
36.7
5524
-47.7%
43.0
2355
-71.8%
40.4
1930
-63.9%
DeepSeek-V4.1-Flash
84.6
1781
-37.7%
68.7
2892
-60.3%
36.0
5949
-43.7%
48.7
3371
-59.6%
38.0
2822
-47.2%
Table 3: Baseline comparison with frontier LLMs as controllers. The reasoning model is DeepSeek-R1-Distill-Qwen-7B.
Figure 4: Action distributions of frontier LLMs as controllers and MetaCtrl across MATH-500 difficulty levels. The reasoning model is DeepSeek-R1-Distill-Qwen-7B.
Figure 5: Action distribution change by MetaCtrl training. The reasoning model is DeepSeek-R1-Distill-Qwen-7B.
Reward Setting
MATH-500
AMC
AIME2024
OlympiadBench
GPQA Diamond
rcorrect
rwrong
λeff
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Vanilla (No Controller)
85.2
2857
73.5
7279
50.0
10570
48.6
8347
32.8
5349
1.5
-1.0
0.5
88.8
1739
-39.1%
76.1
3175
-56.4%
40.7
6261
-40.8%
52.3
3092
-63.0%
43.4
2789
-47.9%
2.0
-1.0
0.5
90.8
1921
-32.8%
77.6
4162
-42.8%
44.0
8106
-23.3%
53.3
3870
-53.6%
37.9
4818
-9.9%
1.0
-0.5
0.5
85.8
1553
-45.6%
71.1
2545
-65.0%
40.7
5013
-52.6%
45.9
2590
-69.0%
36.9
1998
-62.6%
1.0
-1.0
1.0
82.0
1284
-55.1%
59.0
2053
-71.8%
30.7
4626
-56.2%
40.9
1968
-76.4%
37.5
1463
-72.6%
Table 4: Ablation study on reward-function coefficients. The reasoning model is DeepSeek-R1-Distill-Qwen-7B.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Action
Intervention Prompt
Continue
No textual intervention.
Fast Think
Let’s keep only the essential derivation, avoid repetition, and proceed efficiently.
Skip Think
Let’s solve this directly with the minimum necessary reasoning and move to the answer.
Stop Think
The reasoning is sufficient; state only the final conclusion briefly and close the reasoning now.
Appendix
Table 5: Textual intervention prompts associated with the MetaCtrl action space.
Methods
AMC
LiveCodeBench
Acc. ↑
Len. ↓
Reduc. ↑
Acc. ↑
Len. ↓
Reduc. ↑
DeepSeek-R1-Distill-Qwen-7B [Seen Reasoner]
Vanilla
73.5
7279
38.4
10,516
MetaCtrl (ours)
74.7
3036
-58.3%
42.3
2,932
-72.1%
Qwen3-8B [Unseen Reasoner]
Vanilla
78.8
9083
64.7
8,923
Appendix
Table 6: Performance and reasoning efficiency across unseen reasoners on AMC and LiveCodeBench.
Methods
MATH-500
AIME 2024
OlympiadBench
GPQA Diamond
Acc. ↑
Len. ↓
Reduc. ↑
Acc. ↑
Len. ↓
Reduc. ↑
Acc. ↑
Len. ↓
Reduc. ↑
Acc. ↑
Len. ↓
Reduc. ↑
DeepSeek-R1-Distill-Llama-8B [Unseen Reasoner]
Vanilla
87.6
3928
40.0
13472
47.3
8416
26.3
9655
MetaCtrl (ours)
89.0
1942
-50.6%
40.0
7642
-43.3%
50.1
3819
-54.6%
37.4
3192
-66.9%
Appendix
Table 7: Results of DeepSeek-R1-Distill-Llama-8B across five reasoning benchmarks. Acc. denotes pass@1 accuracy, Len. denotes the average total generation length, and Reduc. denotes the relative reduction in generation length compared with vanilla reasoning.
Figure 6: Controller action distributions across different reasoners and benchmarks. Each row corresponds to a reasoner and each column to a benchmark. The learned policy exhibits distinct task–reasoner-specific control patterns rather than a uniform intervention strategy. DeepSeek-R1-7B is used as shorthand for DeepSeek-R1-Distilled-Qwen2.5-7B.
Ablation Setting
MATH-500
AMC
AIME2024
OlympiadBench
GPQA Diamond
Component
Setting
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Vanilla (No Controller)
85.2
2857
73.5
7279
50.0
10570
48.6
8347
32.8
5349
Intervention Position
At Transition Words
71.2
999
-65.0%
47.7
1236
-83.0%
22.7
2828
-73.2%
32.4
1315
-84.2%
35.5
1069
-80.0%
Every 300 Tokens
86.0
3638
+27.3%
68.0
9351
+28.5%
41.3
15500
+46.6%
50.7
9587
+14.9%
29.0
15362
+187.2%
Intervention Action
Add Check Action
87.2
2025
-29.1%
70.8
4000
-45.0%
42.7
8793
-16.8%
51.1
3888
-53.4%
38.2
3746
-30.0%
Controller Model
Initialized from Qwen3-1.7B
88.0
2295
-19.7%
72.5
4850
-33.4%
41.3
10179
-3.7%
51.6
4946
-40.7%
40.9
4600
-14.0%
Appendix
Table 8: Ablation studies on intervention position, action space, and controller initialization. The reasoning model is DeepSeek-R1-Distill-Qwen-7B.
Reasoner
MATH500
AMC
GPQA Diamond
Avg. Time (s)
Mem. Frac. (R+C)
Avg. Time (s)
Mem. Frac. (R+C)
Avg. Time (s)
Mem. Frac. (R+C)
Vanilla Reasoning
Qwen3-8B
31.62
0.90+0.00
74.93
0.90+0.00
81.28
0.90+0.00
Qwen3-14B
45.18
0.90+0.00
100.94
0.90+0.00
95.79
0.90+0.00
MetaCtrl Controlled Reasoning (Two Separate GPUs)
Qwen3-8B
15.44
0.90+0.60
43.02
0.90+0.60
23.96
0.90+0.60
Appendix
Table 9: End-to-end inference latency under Vanilla Reasoning and MetaCtrl Controlled Reasoning with different GPU memory allocations. All experiments are conducted on NVIDIA H20 GPUs. Mem. Frac. (R+C) denotes the configured static GPU memory fractions for the reasoner ( R ) and controller ( C ), respectively. Vanilla Reasoning uses a single GPU for the reasoner. For MetaCtrl with separate GPUs, the reasoner and controller run on two dedicated GPUs. For MetaCtrl with a shared GPU, the reasoner and controller are colocated on the same GPU.
Figure 7: Training dynamics of the MetaCtrl. (a) Training reward during GRPO optimization. (b) Validation accuracy and average total generation tokens throughout training. The reward and accuracy progressively improve and stabilize, while generation length converges to a relatively stable range, indicating that the controller learns to regulate reasoning computation without uniformly compressing reasoning trajectories.
Figure 8: Evolution of the controller action distribution during training. The learned policy gradually shifts toward Continue , while the frequencies of Fast-Think , Skip-Think , and Stop-Think decrease and stabilize at lower levels. This trend suggests that the controller progressively learns to intervene more selectively rather than frequently modifying the reasoner’s trajectory.
Figure 9: Vanilla reasoning on a MATH-500 example. The model obtains the correct answer but continues with repetitive verification and alternative derivations, leading to substantial overthinking and 7,043 reasoning tokens.
Figure 10: MetaCtrl reasoning on the MATH-500 example. Both methods arrive at the correct answer, while MetaCtrl dynamically regulates the reasoning trajectory through lightweight interventions and terminates once the solution is sufficiently established, reducing the reasoning length from 7,043 to 957 tokens.
Figure 11: Vanilla reasoning on a MATH-500 example. The model reaches the correct answer early but continues with alternative counting strategies, repeated verification, and additional sanity checks, resulting in 2,479 reasoning tokens.
Figure 12: MetaCtrl reasoning on the MATH-500 example. After identifying that all valid handshakes occur between the two groups, the controller skips unnecessary intermediate reasoning and terminates once sufficient information for the solution has been established, reducing the reasoning length from 2,479 to 409 tokens while preserving the correct answer.
Figure 13: Vanilla reasoning on a GPQA example. Although the model eventually selects the correct answer, it repeatedly revisits competing explanations and re-evaluates previously considered options, resulting in 2,305 reasoning tokens.
Figure 14: MetaCtrl reasoning on the same GPQA example. The controller progressively regulates the reasoning trajectory through Fast Think, Skip Think, and Stop Think interventions, reducing repeated hypothesis exploration and reaching the same correct answer with 707 reasoning tokens.
Figure 15: Vanilla reasoning on an AIME 2024 example. The model derives the correct composition-based counting argument but continues to revisit the same calculation through repeated verification and alternative checks, resulting in 6,637 reasoning tokens.
Figure 16: MetaCtrl reasoning on the same AIME 2024 example. The controller dynamically accelerates the reasoning trajectory and terminates once the composition-based counting argument is complete, reaching the same correct answer with 1,789 reasoning tokens.
Figure 17: Vanilla reasoning on an OlympiadBench example. The model correctly derives the unique solution but continues with repeated substitution checks, alternative derivations, and unnecessary uniqueness verification, resulting in 6,064 reasoning tokens.
Figure 18: MetaCtrl reasoning on the same OlympiadBench example. After the key algebraic reduction establishes the unique solution, the controller terminates further deliberation, reaching the same correct answer with 1,463 reasoning tokens.
Figure 19: Vanilla reasoning on an AMC example. The model enters an inconsistent trigonometric detour and continues reasoning despite recognizing contradictions in its intermediate derivation, ultimately producing an incorrect answer after 8,535 reasoning tokens.
Figure 20: MetaCtrl reasoning on the same AMC example. By dynamically regulating how the frozen reasoner proceeds, MetaCtrl follows a more direct derivation and avoids prolonged unproductive deliberation, reaching the correct answer with 2,479 reasoning tokens.
Figure 21: Vanilla reasoning on a GPQA chemistry example. The model repeatedly revisits conflicting hypotheses about solvation, basicity, and nucleophilicity, and an early mischaracterization of the alkoxide contributes to a prolonged inconsistent trajectory, ultimately yielding an incorrect answer after 12,806 reasoning tokens.
Figure 22: MetaCtrl reasoning on the same GPQA example. By dynamically regulating how the frozen reasoner proceeds, MetaCtrl substantially reduces prolonged hypothesis exploration and reaches the correct answer with 2,001 reasoning tokens.
Figure 23: Vanilla reasoning on a GPQA physics example. The model repeatedly interprets the “first two minima” as the symmetric first minima on opposite sides of the central maximum, leading to prolonged verification of the same interpretation and ultimately an incorrect answer after 3,702 reasoning tokens.
Figure 24: MetaCtrl reasoning on the same GPQA example. The controlled trajectory revisits the interpretation of the diffraction minima and uses the first two zeros of the Airy pattern, reaching the correct answer with 3,117 reasoning tokens.
Large Reasoning Models (LRMs) can exhibit step-by-step reasoning, reflection, and backtracking, but these behaviors are often unregulated, leading to overthinking. As a result, LRMs continue generating redundant reasoning even after reaching high-confidence conclusions. This increases inference cost and latency, limiting practical deployment. The root cause is the absence of an intrinsic mechanism to monitor the reasoning state and decide when to continue, backtrack, or stop. We propose MERA, a meta-cognitive reasoning framework that decouples reasoning from control to enable independent optimization of control strategies. MERA constructs high-quality reasoning-control supervision data via a takeover-based pipeline, and transforms long-horizon traces into structured reasoning-control alternating sequences for training. The model is trained with supervised fine-tuning to internalize the structured separation, and further optimized with Control-Segment Policy Optimization (CSPO), which combines segment-wise GRPO with control masking to focus learning on control segments. Experiments across reasoning benchmarks show that MERA improves both efficiency and accuracy.
Rui Ha, Rui Pu, Chaozhuo Li +2
Beijing University of Posts and Telecommunications, China · 2Chongqing University of Posts and Telecommunications, China
Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reasoning methods control thinking length by shortening, early-stopping, or compressing traces, leaving how the model thinks implicit. In this paper, we propose Agentic Chain-of-Thought Steering (ACTS), which formulates reasoning steering as a Markov decision process where a controller agent adaptively steers a frozen reasoner during inference. At each step, the controller observes the reasoning trace and remaining thinking budget, then issues a steering action consisting of a reasoning strategy and a steering phrase that initiates the next reasoner step. This enables budget-aware strategy control for efficient reasoning while preserving the reasoner's generation continuity. We initialize the controller agent from our constructed synthetic steering trajectories with multi-budget augmentation, and further optimize it via reinforcement learning with budget-conditioned reward shaping. Experiments across multiple benchmarks show that ACTS matches full-thinking performance with substantial token savings, and enables controllable accuracy-efficiency trade-offs across different reasoners and tasks. The code is available at https://github.com/Andree-9/ACTS.
Yu Xia, Zhouhang Xie, Xin Xu +4
University of California San Diego1 · Intuit AI Research2
Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt to mitigate this issue but inevitably truncate necessary steps and induce underthinking, thereby compromising performance. To address this dilemma, we investigate the latent representations and observe that efficient reasoning steps naturally cluster into a concentrated region in latent space, while those deviating from this region tend to produce verbose sequences. To leverage this, we keep reasoning focused within this region via a quadratic program which projects deviating hidden states back into the region. Then we propose a novel training-free framework to achieve efficient reasoning that reduces token generation costs without sacrificing performance. Extensive experiments conducted on four models ranging from 1.5B to 14B, and across six benchmarks in math reasoning, coding, and scientific QA, validate the effectiveness of our method, up to a 12.1% improvement in accuracy while reducing generated tokens by 11.8% to 52.8%. Codes are available at \href{https://github.com/hzn18/Opt4Reasoning}{https://github.com/hzn18/Opt4Reasoning}.