Organizations: Interdisciplinary Center for Security, Reliability, and Trust, University of Luxembourg, Luxembourg · Department of Social and Decision Sciences, Carnegie Mellon University, Pittsburgh PA
Metacognition involves reasoning about cognitive processes themselves. An example is in resource allocation where we choose how much time and effort to put into a reasoning task before we begin based on our confidence. Current Artificial Intelligence (AI) systems that rely on Large Language Models (LLMs) cannot estimate their uncertainty about an output without first responding, and cannot dynamically allocate resources to producing an output, making this type of metacognitive process difficult. A recently proposed alternative to classic transformer architectures that addresses these two concerns is the Energy Based Model (EBM) which allows for interpretable uncertainty modeling and dynamic allocation of compute resources. While EBMs can allow for control of these two processes, the actual metacognitive task of determining compute allocation based on uncertainty is not directly addressed. Instance-Based Learning Theory (IBLT) provides an approach to modeling human-like decisions from experience that has previously been applied to predicting human metacognitive reasoning. In this paper we introduce a framework for MEtacognitive Reasoning with Instance-based Learning Theory and Energy Dynamics (MERITED). Grounded in IBLT, this framework allows for control of the computational effort allocated in an EBM to allow for metacognitive control over reasoning effort based on uncertainty while remaining computationally efficient. This work has two main contributions, the training and open weight sharing of a 191M parameter reasoning EBM, and an implementation of the MERITED framework for dynamic compute allocation using an IBL model.
Figures & tables
Figure 1: Three examples of Generative Models. Previous applications of IBL onto GMs have integrated left: Generative Adversarial Networks (GANs) and middle: Auto-regressive Transformers (ATs). In this work we integrate right: an Energy-Based Model (EBM). Each sub-figure is structured as an input-process-output model, with blue boxes being the input, green boxes being the process, and the orange and black boxes being the output. EBT has an additional step based on the energy function output that involves optimization. The orange box in each sub-figure is the representation that is used to calculate similarity for IBL models of the complex stimuli in a decision making task.
Figure 2: Orange boxes represent the embedding and energy distribution used by the IBL model. Here, the decision made by the IBL model is either what depth to use (step selection) or whether to stop optimization (optimal stopping). The EBT generates the energy distribution and question embeddings fed into IBL. The final framework output from processing the data and candidate tokens after the optimization step is a probability distribution.
Figure 3: Left Panel: Question response accuracy for the ARC-Easy (blue) and OpenBookQA (green) datasets. Right Panel: Orange bars indicate the average runtime for answering the question across both tasks in seconds. The dashed line represents the highest accuracy by selecting from all EBT depths for the ARC dataset, and the dotted line for the OpenBookQA dataset. Random choice performance for all questions is 25%.
Model
ARC (%)
OB (%)
Std
Time (s)
EBT d=1
52.00
26.40
18.10
0.35
EBT d=2
55.00 ∗
26.60
20.08
0.65
EBT d=4
53.00
25.60
19.37
1.25
EBT d=8
51.00
27.20
16.83
2.45
EBT d=10
55.00 ∗
26.60
20.08
3.07
MERITED Depth
58.00 ∗
29.40
20.22
2.09
Table 1: ARC accuracy, OB accuracy, Accuracy standard deviation, and runtime for each of the 5 static MCMC depth models, the two MERITED methods, and the GPT-2 baseline performance. Omniscient performance is theoretical limit of MERITED performance.
Component
Configuration
Base model
Same pretrained GPT-2 small checkpoint as Table 3
Level 1
Fast direct GPT-2 candidate scoring
Routing signal
Difference between the two largest softmax-normalized candidate scores
Confidence threshold
0.50 margin; lower-confidence questions are routed for additional reasoning
Routed reasoning
Level 2 situational-awareness prompt containing the question and candidate choices
Final answer
Candidate with the highest rescored GPT-2 log-probability
Table 2: GPT-2-CogRouter evaluation configuration. This table describes the router wrapper used in the comparison, not a separately fine-tuned GPT-2 checkpoint.
EBT-191M
GPT-2 small (SLM)
Architecture
Parameters
∼ 191M
124M
Transformer blocks
12
12
Embedding dimension
1024
768
Attention heads
16
12
Head dimension
64
64
Table 3: Model architecture and training configuration for the 191M Energy-Based Transformer (EBT) and the GPT-2 small baseline (SLM) used in the ARC-Easy multiple-choice comparison. Dashes (—) denote parameters that do not apply to the standard autoregressive baseline.
Model
ARC Acc (%)
OB Acc (%)
Mean Acc (%)
Acc Std
Avg Runtime (s)
Runtime Std
EBT d=1
52.00
26.40
39.20
18.10
0.35
0.02
EBT d=2
55.00 ∗
26.60
40.80
20.08
0.65
0.07
EBT d=4
53.00
25.60
39.30
19.37
1.25
0.15
EBT d=8
51.00
27.20
39.10
16.83
2.45
0.29
EBT d=10
55.00 ∗
26.60
40.80
20.08
3.07
0.39
MERITED depth-choice
58.00 ∗
29.40
43.70 ∗
20.22
2.09
0.35
Table 4: Combined two-task best-sweep comparison across ARC-Easy and OBQA. GPT-2 is the baseline row. Asterisk marks one-sided significance over GPT-2 baseline at α=0.05 . Acc Std and Runtime Std are computed across the two task-level values for each model. Bolding represents best performance within column, ignoring the hypothetical optimal depth-choice.
Large Reasoning Models (LRMs) can exhibit step-by-step reasoning, reflection, and backtracking, but these behaviors are often unregulated, leading to overthinking. As a result, LRMs continue generating redundant reasoning even after reaching high-confidence conclusions. This increases inference cost and latency, limiting practical deployment. The root cause is the absence of an intrinsic mechanism to monitor the reasoning state and decide when to continue, backtrack, or stop. We propose MERA, a meta-cognitive reasoning framework that decouples reasoning from control to enable independent optimization of control strategies. MERA constructs high-quality reasoning-control supervision data via a takeover-based pipeline, and transforms long-horizon traces into structured reasoning-control alternating sequences for training. The model is trained with supervised fine-tuning to internalize the structured separation, and further optimized with Control-Segment Policy Optimization (CSPO), which combines segment-wise GRPO with control masking to focus learning on control segments. Experiments across reasoning benchmarks show that MERA improves both efficiency and accuracy.
Rui Ha, Rui Pu, Chaozhuo Li +2
Beijing University of Posts and Telecommunications, China · 2Chongqing University of Posts and Telecommunications, China
Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Conversely, aggressively shortening reasoning can degrade performance on difficult ones. Effective reasoning therefore requires dynamically deciding when additional computation is useful based on the reasoner's capabilities and evolving solution state. Existing approaches often rely on predefined budgets or intervention rules, retrain the target reasoner, or require additional supervision. We introduce MetaCtrl, a lightweight controller that adaptively regulates a frozen reasoner without predefined token budgets or reasoner retraining. We formulate reasoning regulation as a sequential metacognitive control problem: MetaCtrl observes the evolving reasoning trace and decides whether to continue, simplify, skip redundant steps, or conclude reasoning. It is trained directly with reinforcement learning using a reward that prioritizes correctness while favoring shorter trajectories among correct solutions, requiring neither supervised intervention trajectories nor problem-specific budgets. Across seven benchmarks spanning mathematics, science, and code, MetaCtrl consistently improves the accuracy of LRMs while reducing their reasoning length. On DeepSeek-R1-Distill-Qwen-7B, it improves average accuracy by 4.7 points while reducing generation length by 53.3%. Without further training, the same controller transfers to an unseen reasoner (e.g., Qwen3-14B), improving average accuracy by 2.9 points and reducing generation length by 50.3%. These results establish MetaCtrl as a plug-and-play controller for improving reasoning accuracy while substantially reducing inference-time generation. The code is available at https://github.com/binbin2xs/MetaCtrl.
Zhibin Wen, Tao Han, Lei Bai +2
Southern University of Science and Technology · The Hong Kong University of Science and Technology · Shanghai Artificial Intelligence Laboratory +1
Large Reasoning Models (LRMs) generate chain-of-thought traces whose length tracks human reaction times across cognitive tasks, but recent debate questions whether this alignment reflects genuine computational structure or surface verbosity. We test whether the alignment varies with inference-time reasoning effort. Across GPT-OSS-20B and GPT-OSS-120B, three effort levels, and six reasoning tasks, within-task and cross-task alignment remain invariant: Bayes Factors lean toward the null, and mean alignment is numerically near-identical across conditions. A manipulation check reveals that the effort parameter sets an upper budget on generation rather than driving real-time allocation, suggesting that the allocation policy is crystallized at training time. Arithmetic complexity contrasts further show that token allocation tracks fine-grained, format-dependent human difficulty patterns, with model scale improving the match. Cognitive cost alignment between LRMs and humans appears to be a training-time achievement, robust to inference-time perturbations, supporting a compiled rather than online account of LRM problem-solving.
Yueqing Hu, Tianhong Wang
Institute of Neuroscience, Chinese Academy of Sciences, Shanghai, China · School of Philosophy, Anhui University, Hefei, China