MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller
Authors: Zhibin Wen, Tao Han, Lei Bai, Can Li, Yang Xu
Organizations: Southern University of Science and Technology · The Hong Kong University of Science and Technology · Shanghai Artificial Intelligence Laboratory · Tongji University
Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Conversely, aggressively shortening reasoning can degrade performance on difficult ones. Effective reasoning therefore requires dynamically deciding when additional computation is useful based on the reasoner's capabilities and evolving solution state. Existing approaches often rely on predefined budgets or intervention rules, retrain the target reasoner, or require additional supervision. We introduce MetaCtrl, a lightweight controller that adaptively regulates a frozen reasoner without predefined token budgets or reasoner retraining. We formulate reasoning regulation as a sequential metacognitive control problem: MetaCtrl observes the evolving reasoning trace and decides whether to continue, simplify, skip redundant steps, or conclude reasoning. It is trained directly with reinforcement learning using a reward that prioritizes correctness while favoring shorter trajectories among correct solutions, requiring neither supervised intervention trajectories nor problem-specific budgets. Across seven benchmarks spanning mathematics, science, and code, MetaCtrl consistently improves the accuracy of LRMs while reducing their reasoning length. On DeepSeek-R1-Distill-Qwen-7B, it improves average accuracy by 4.7 points while reducing generation length by 53.3%. Without further training, the same controller transfers to an unseen reasoner (e.g., Qwen3-14B), improving average accuracy by 2.9 points and reducing generation length by 50.3%. These results establish MetaCtrl as a plug-and-play controller for improving reasoning accuracy while substantially reducing inference-time generation. The code is available at https://github.com/binbin2xs/MetaCtrl.
Figures & tables
Figure 1: Left: Accuracy and generation length of the same reasoner when controlled by MetaCtrl, controlled by frontier LLMs, or run without external control across five benchmarks. Right: Comparison of MetaCtrl with representative efficient reasoning baselines across two reasoning models.
Figure 2: Comparison of four paradigms for efficient LRMs inference. In external control methods, some approaches train the reasoner, while others keep it frozen.
Figure 3: Overview of MetaCtrl. (a) A learned meta-level controller regulates a frozen reasoner through four intervention actions, without an input budget. (b) At each turn, it selects an intervention based on the current reasoning state. (c) The controller is trained with GRPO using an efficiency-aware outcome reward. (d) The trained controller generalizes to unseen reasoners without retraining.
Math Reasoning
Scientific Reasoning
Methods
MATH-500
AIME2024
OmniMath
GPQA Diamond
AVG
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc. ↑
Reduc. ↑
DeepSeek-R1-Distill-Qwen-7B [Seen Reasoner]
NoThinking
79.4
699
-75.5
40.0
4134
-60.9
40.0
2083
-63.7
29.3
2084
-61.0
47.2
-65.3
Vanilla
85.2
2857
−0%
50.0
10570
−0%
45.0
5736
−0%
32.8
5349
−0%
53.3
−0%
D-Prompt
86.0
2726
−4.6%
46.7
10209
−3.4%
45.0
5398
−5.9%
30.3
5117
−4.3%
52.0
−4.6%
Table 1: Comparison across reasoning models of different scales and and multiple baselines.
Math Reasoning
Scientific Reasoning
Methods
MATH-500
AIME 2024
OlympiadBench
GPQA Diamond
AVG
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc. ↑
Reduc. ↑
Qwen3-8B [Unseen Reasoner]
NoThinking
87.4
1480
-70.0
23.3
7121
-41.2
48.7
5219
-43.7
48.5
2564
-72.7
52.0
-56.9
Vanilla
92.2
4926
−0%
63.3
12101
−0%
59.9
9268
−0%
52.5
9382
−0%
67.0
−0%
TALE
91.8
3682
−25.3%
60.0
11847
−2.1%
56.1
7306
−21.2%
52.0
4161
−55.6%
65.0
−26.0%
Table 2: Comparison across the Qwen3 series at different model scales and multiple baselines.
Controller Model
MATH-500
AMC
AIME2024
OlympiadBench
GPQA Diamond
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Vanilla (No Controller)
85.2
2857
73.5
7279
50.0
10570
48.6
8347
32.8
5349
Qwen-3.8-Flash
71.4
6548
+129.2%
60.2
9229
+26.8%
–
–
–
–
–
–
30.3
10584
+97.9%
Qwen-3.8-Max
78.2
1605
-43.8%
48.8
2507
-65.6%
30.0
7780
-26.4%
39.3
2921
-65.0%
38.4
1813
-66.1%
GLM-5.2
82.0
1472
-48.5%
66.9
2450
-66.3%
36.7
5524
-47.7%
43.0
2355
-71.8%
40.4
1930
-63.9%
DeepSeek-V4.1-Flash
84.6
1781
-37.7%
68.7
2892
-60.3%
36.0
5949
-43.7%
48.7
3371
-59.6%
38.0
2822
-47.2%
Table 3: Baseline comparison with frontier LLMs as controllers. The reasoning model is DeepSeek-R1-Distill-Qwen-7B.
Figure 4: Action distributions of frontier LLMs as controllers and MetaCtrl across MATH-500 difficulty levels. The reasoning model is DeepSeek-R1-Distill-Qwen-7B.
Figure 5: Action distribution change by MetaCtrl training. The reasoning model is DeepSeek-R1-Distill-Qwen-7B.
Reward Setting
MATH-500
AMC
AIME2024
OlympiadBench
GPQA Diamond
rcorrect
rwrong
λeff
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Vanilla (No Controller)
85.2
2857
73.5
7279
50.0
10570
48.6
8347
32.8
5349
1.5
-1.0
0.5
88.8
1739
-39.1%
76.1
3175
-56.4%
40.7
6261
-40.8%
52.3
3092
-63.0%
43.4
2789
-47.9%
2.0
-1.0
0.5
90.8
1921
-32.8%
77.6
4162
-42.8%
44.0
8106
-23.3%
53.3
3870
-53.6%
37.9
4818
-9.9%
1.0
-0.5
0.5
85.8
1553
-45.6%
71.1
2545
-65.0%
40.7
5013
-52.6%
45.9
2590
-69.0%
36.9
1998
-62.6%
1.0
-1.0
1.0
82.0
1284
-55.1%
59.0
2053
-71.8%
30.7
4626
-56.2%
40.9
1968
-76.4%
37.5
1463
-72.6%
Table 4: Ablation study on reward-function coefficients. The reasoning model is DeepSeek-R1-Distill-Qwen-7B.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Action
Intervention Prompt
Continue
No textual intervention.
Fast Think
Let’s keep only the essential derivation, avoid repetition, and proceed efficiently.
Skip Think
Let’s solve this directly with the minimum necessary reasoning and move to the answer.
Stop Think
The reasoning is sufficient; state only the final conclusion briefly and close the reasoning now.
Appendix
Table 5: Textual intervention prompts associated with the MetaCtrl action space.
Methods
AMC
LiveCodeBench
Acc. ↑
Len. ↓
Reduc. ↑
Acc. ↑
Len. ↓
Reduc. ↑
DeepSeek-R1-Distill-Qwen-7B [Seen Reasoner]
Vanilla
73.5
7279
38.4
10,516
MetaCtrl (ours)
74.7
3036
-58.3%
42.3
2,932
-72.1%
Qwen3-8B [Unseen Reasoner]
Vanilla
78.8
9083
64.7
8,923
Appendix
Table 6: Performance and reasoning efficiency across unseen reasoners on AMC and LiveCodeBench.
Methods
MATH-500
AIME 2024
OlympiadBench
GPQA Diamond
Acc. ↑
Len. ↓
Reduc. ↑
Acc. ↑
Len. ↓
Reduc. ↑
Acc. ↑
Len. ↓
Reduc. ↑
Acc. ↑
Len. ↓
Reduc. ↑
DeepSeek-R1-Distill-Llama-8B [Unseen Reasoner]
Vanilla
87.6
3928
40.0
13472
47.3
8416
26.3
9655
MetaCtrl (ours)
89.0
1942
-50.6%
40.0
7642
-43.3%
50.1
3819
-54.6%
37.4
3192
-66.9%
Appendix
Table 7: Results of DeepSeek-R1-Distill-Llama-8B across five reasoning benchmarks. Acc. denotes pass@1 accuracy, Len. denotes the average total generation length, and Reduc. denotes the relative reduction in generation length compared with vanilla reasoning.
Figure 6: Controller action distributions across different reasoners and benchmarks. Each row corresponds to a reasoner and each column to a benchmark. The learned policy exhibits distinct task–reasoner-specific control patterns rather than a uniform intervention strategy. DeepSeek-R1-7B is used as shorthand for DeepSeek-R1-Distilled-Qwen2.5-7B.
Ablation Setting
MATH-500
AMC
AIME2024
OlympiadBench
GPQA Diamond
Component
Setting
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Acc.
Len.
Reduc.
Vanilla (No Controller)
85.2
2857
73.5
7279
50.0
10570
48.6
8347
32.8
5349
Intervention Position
At Transition Words
71.2
999
-65.0%
47.7
1236
-83.0%
22.7
2828
-73.2%
32.4
1315
-84.2%
35.5
1069
-80.0%
Every 300 Tokens
86.0
3638
+27.3%
68.0
9351
+28.5%
41.3
15500
+46.6%
50.7
9587
+14.9%
29.0
15362
+187.2%
Intervention Action
Add Check Action
87.2
2025
-29.1%
70.8
4000
-45.0%
42.7
8793
-16.8%
51.1
3888
-53.4%
38.2
3746
-30.0%
Controller Model
Initialized from Qwen3-1.7B
88.0
2295
-19.7%
72.5
4850
-33.4%
41.3
10179
-3.7%
51.6
4946
-40.7%
40.9
4600
-14.0%
Appendix
Table 8: Ablation studies on intervention position, action space, and controller initialization. The reasoning model is DeepSeek-R1-Distill-Qwen-7B.
Reasoner
MATH500
AMC
GPQA Diamond
Avg. Time (s)
Mem. Frac. (R+C)
Avg. Time (s)
Mem. Frac. (R+C)
Avg. Time (s)
Mem. Frac. (R+C)
Vanilla Reasoning
Qwen3-8B
31.62
0.90+0.00
74.93
0.90+0.00
81.28
0.90+0.00
Qwen3-14B
45.18
0.90+0.00
100.94
0.90+0.00
95.79
0.90+0.00
MetaCtrl Controlled Reasoning (Two Separate GPUs)
Qwen3-8B
15.44
0.90+0.60
43.02
0.90+0.60
23.96
0.90+0.60
Appendix
Table 9: End-to-end inference latency under Vanilla Reasoning and MetaCtrl Controlled Reasoning with different GPU memory allocations. All experiments are conducted on NVIDIA H20 GPUs. Mem. Frac. (R+C) denotes the configured static GPU memory fractions for the reasoner ( R ) and controller ( C ), respectively. Vanilla Reasoning uses a single GPU for the reasoner. For MetaCtrl with separate GPUs, the reasoner and controller run on two dedicated GPUs. For MetaCtrl with a shared GPU, the reasoner and controller are colocated on the same GPU.
Figure 7: Training dynamics of the MetaCtrl. (a) Training reward during GRPO optimization. (b) Validation accuracy and average total generation tokens throughout training. The reward and accuracy progressively improve and stabilize, while generation length converges to a relatively stable range, indicating that the controller learns to regulate reasoning computation without uniformly compressing reasoning trajectories.
Figure 8: Evolution of the controller action distribution during training. The learned policy gradually shifts toward Continue , while the frequencies of Fast-Think , Skip-Think , and Stop-Think decrease and stabilize at lower levels. This trend suggests that the controller progressively learns to intervene more selectively rather than frequently modifying the reasoner’s trajectory.
Figure 9: Vanilla reasoning on a MATH-500 example. The model obtains the correct answer but continues with repetitive verification and alternative derivations, leading to substantial overthinking and 7,043 reasoning tokens.
Figure 10: MetaCtrl reasoning on the MATH-500 example. Both methods arrive at the correct answer, while MetaCtrl dynamically regulates the reasoning trajectory through lightweight interventions and terminates once the solution is sufficiently established, reducing the reasoning length from 7,043 to 957 tokens.
Figure 11: Vanilla reasoning on a MATH-500 example. The model reaches the correct answer early but continues with alternative counting strategies, repeated verification, and additional sanity checks, resulting in 2,479 reasoning tokens.
Figure 12: MetaCtrl reasoning on the MATH-500 example. After identifying that all valid handshakes occur between the two groups, the controller skips unnecessary intermediate reasoning and terminates once sufficient information for the solution has been established, reducing the reasoning length from 2,479 to 409 tokens while preserving the correct answer.
Figure 13: Vanilla reasoning on a GPQA example. Although the model eventually selects the correct answer, it repeatedly revisits competing explanations and re-evaluates previously considered options, resulting in 2,305 reasoning tokens.
Figure 14: MetaCtrl reasoning on the same GPQA example. The controller progressively regulates the reasoning trajectory through Fast Think, Skip Think, and Stop Think interventions, reducing repeated hypothesis exploration and reaching the same correct answer with 707 reasoning tokens.
Figure 15: Vanilla reasoning on an AIME 2024 example. The model derives the correct composition-based counting argument but continues to revisit the same calculation through repeated verification and alternative checks, resulting in 6,637 reasoning tokens.
Figure 16: MetaCtrl reasoning on the same AIME 2024 example. The controller dynamically accelerates the reasoning trajectory and terminates once the composition-based counting argument is complete, reaching the same correct answer with 1,789 reasoning tokens.
Figure 17: Vanilla reasoning on an OlympiadBench example. The model correctly derives the unique solution but continues with repeated substitution checks, alternative derivations, and unnecessary uniqueness verification, resulting in 6,064 reasoning tokens.
Figure 18: MetaCtrl reasoning on the same OlympiadBench example. After the key algebraic reduction establishes the unique solution, the controller terminates further deliberation, reaching the same correct answer with 1,463 reasoning tokens.
Figure 19: Vanilla reasoning on an AMC example. The model enters an inconsistent trigonometric detour and continues reasoning despite recognizing contradictions in its intermediate derivation, ultimately producing an incorrect answer after 8,535 reasoning tokens.
Figure 20: MetaCtrl reasoning on the same AMC example. By dynamically regulating how the frozen reasoner proceeds, MetaCtrl follows a more direct derivation and avoids prolonged unproductive deliberation, reaching the correct answer with 2,479 reasoning tokens.
Figure 21: Vanilla reasoning on a GPQA chemistry example. The model repeatedly revisits conflicting hypotheses about solvation, basicity, and nucleophilicity, and an early mischaracterization of the alkoxide contributes to a prolonged inconsistent trajectory, ultimately yielding an incorrect answer after 12,806 reasoning tokens.
Figure 22: MetaCtrl reasoning on the same GPQA example. By dynamically regulating how the frozen reasoner proceeds, MetaCtrl substantially reduces prolonged hypothesis exploration and reaches the correct answer with 2,001 reasoning tokens.
Figure 23: Vanilla reasoning on a GPQA physics example. The model repeatedly interprets the “first two minima” as the symmetric first minima on opposite sides of the central maximum, leading to prolonged verification of the same interpretation and ultimately an incorrect answer after 3,702 reasoning tokens.
Figure 24: MetaCtrl reasoning on the same GPQA example. The controlled trajectory revisits the interpretation of the diffraction minima and uses the first two zeros of the Airy pattern, reaching the correct answer with 3,117 reasoning tokens.