Organizations: Interdisciplinary Center for Security, Reliability, and Trust, University of Luxembourg, Luxembourg · Department of Social and Decision Sciences, Carnegie Mellon University, Pittsburgh PA
Metacognition involves reasoning about cognitive processes themselves. An example is in resource allocation where we choose how much time and effort to put into a reasoning task before we begin based on our confidence. Current Artificial Intelligence (AI) systems that rely on Large Language Models (LLMs) cannot estimate their uncertainty about an output without first responding, and cannot dynamically allocate resources to producing an output, making this type of metacognitive process difficult. A recently proposed alternative to classic transformer architectures that addresses these two concerns is the Energy Based Model (EBM) which allows for interpretable uncertainty modeling and dynamic allocation of compute resources. While EBMs can allow for control of these two processes, the actual metacognitive task of determining compute allocation based on uncertainty is not directly addressed. Instance-Based Learning Theory (IBLT) provides an approach to modeling human-like decisions from experience that has previously been applied to predicting human metacognitive reasoning. In this paper we introduce a framework for MEtacognitive Reasoning with Instance-based Learning Theory and Energy Dynamics (MERITED). Grounded in IBLT, this framework allows for control of the computational effort allocated in an EBM to allow for metacognitive control over reasoning effort based on uncertainty while remaining computationally efficient. This work has two main contributions, the training and open weight sharing of a 191M parameter reasoning EBM, and an implementation of the MERITED framework for dynamic compute allocation using an IBL model.
Figures & tables
Figure 1: Three examples of Generative Models. Previous applications of IBL onto GMs have integrated left: Generative Adversarial Networks (GANs) and middle: Auto-regressive Transformers (ATs). In this work we integrate right: an Energy-Based Model (EBM). Each sub-figure is structured as an input-process-output model, with blue boxes being the input, green boxes being the process, and the orange and black boxes being the output. EBT has an additional step based on the energy function output that involves optimization. The orange box in each sub-figure is the representation that is used to calculate similarity for IBL models of the complex stimuli in a decision making task.
Figure 2: Orange boxes represent the embedding and energy distribution used by the IBL model. Here, the decision made by the IBL model is either what depth to use (step selection) or whether to stop optimization (optimal stopping). The EBT generates the energy distribution and question embeddings fed into IBL. The final framework output from processing the data and candidate tokens after the optimization step is a probability distribution.
Figure 3: Left Panel: Question response accuracy for the ARC-Easy (blue) and OpenBookQA (green) datasets. Right Panel: Orange bars indicate the average runtime for answering the question across both tasks in seconds. The dashed line represents the highest accuracy by selecting from all EBT depths for the ARC dataset, and the dotted line for the OpenBookQA dataset. Random choice performance for all questions is 25%.
Model
ARC (%)
OB (%)
Std
Time (s)
EBT d=1
52.00
26.40
18.10
0.35
EBT d=2
55.00 ∗
26.60
20.08
0.65
EBT d=4
53.00
25.60
19.37
1.25
EBT d=8
51.00
27.20
16.83
2.45
EBT d=10
55.00 ∗
26.60
20.08
3.07
MERITED Depth
58.00 ∗
29.40
20.22
2.09
Table 1: ARC accuracy, OB accuracy, Accuracy standard deviation, and runtime for each of the 5 static MCMC depth models, the two MERITED methods, and the GPT-2 baseline performance. Omniscient performance is theoretical limit of MERITED performance.
Component
Configuration
Base model
Same pretrained GPT-2 small checkpoint as Table 3
Level 1
Fast direct GPT-2 candidate scoring
Routing signal
Difference between the two largest softmax-normalized candidate scores
Confidence threshold
0.50 margin; lower-confidence questions are routed for additional reasoning
Routed reasoning
Level 2 situational-awareness prompt containing the question and candidate choices
Final answer
Candidate with the highest rescored GPT-2 log-probability
Table 2: GPT-2-CogRouter evaluation configuration. This table describes the router wrapper used in the comparison, not a separately fine-tuned GPT-2 checkpoint.
EBT-191M
GPT-2 small (SLM)
Architecture
Parameters
∼ 191M
124M
Transformer blocks
12
12
Embedding dimension
1024
768
Attention heads
16
12
Head dimension
64
64
Table 3: Model architecture and training configuration for the 191M Energy-Based Transformer (EBT) and the GPT-2 small baseline (SLM) used in the ARC-Easy multiple-choice comparison. Dashes (—) denote parameters that do not apply to the standard autoregressive baseline.
Model
ARC Acc (%)
OB Acc (%)
Mean Acc (%)
Acc Std
Avg Runtime (s)
Runtime Std
EBT d=1
52.00
26.40
39.20
18.10
0.35
0.02
EBT d=2
55.00 ∗
26.60
40.80
20.08
0.65
0.07
EBT d=4
53.00
25.60
39.30
19.37
1.25
0.15
EBT d=8
51.00
27.20
39.10
16.83
2.45
0.29
EBT d=10
55.00 ∗
26.60
40.80
20.08
3.07
0.39
MERITED depth-choice
58.00 ∗
29.40
43.70 ∗
20.22
2.09
0.35
Table 4: Combined two-task best-sweep comparison across ARC-Easy and OBQA. GPT-2 is the baseline row. Asterisk marks one-sided significance over GPT-2 baseline at α=0.05 . Acc Std and Runtime Std are computed across the two task-level values for each model. Bolding represents best performance within column, ignoring the hypothetical optimal depth-choice.