Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, and that fine-tuning can substantially improve them. However, what models actually learn during such training remains poorly understood. We investigate how LLMs acquire metacognitive monitoring, the ability to know what one knows, by training 10 open-weight LLMs to predict their own accuracy on factual multiple-choice questions before answering them. We find that trained confidence reflects two distinct signals. While on questions close to the training data, it tracks the model's true accuracy, in other domains, it instead tracks output consistency: the concentration of the model's answer distribution. Output consistency tracking emerges early in training and generalizes across datasets, whereas accuracy tracking develops later and remains local to the training distribution. These results suggest that calibration training may not teach models to generally detect errors they commit confidently, and they raise broader questions about the nature of metacognition in artificial systems.
Figures & tables
Figure 1: Trained confidence tracks accuracy locally and consistency globally Left side: experimental setup — models answer factual multiple-choice questions from five dataset variants. Accuracy is the probability assigned to the correct option, and output consistency is the normalized concentration of the answer distribution; the two can differ sharply, e.g., when a model reliably selects a wrong answer. Before training, verbalized confidence is essentially uncorrelated with true accuracy ( r=0.04 ). After fine-tuning with LoRA and a linear probe, trained confidence predicts accuracy much better ( r=0.73 ). Right side: schematics of the results — trained confidence tracks true accuracy on questions near the training distribution but output consistency on more distant ones, such as held-out topics or other datasets.
Figure 2: Relationship between predicted accuracy, true accuracy and output consistency. Left panel shows the trained confidence of Qwen3.5-9B against true accuracy (top) and output consistency (bottom) on the test split of MedMCQA 4 and Math 4 . Red lines are quadratic fits ( R2 shown above each plot). Right panel represents Δr on each test split; positive values (red) indicate that confidence tracks true accuracy more than output consistency, negative values (blue) the reverse. Error bars: s.e.m. across the 10 models (each averaged over 5 runs).
Figure 3: Cross-dataset evaluation of trained confidence. Each panel shows models trained on one dataset and evaluated on the test split of all five datasets. Bars represent Δr ; positive values (red) indicate that confidence tracks true accuracy more than output consistency, negative values (blue) the reverse. In-distribution evaluations (tested on the test split from the training dataset - thematically distinct but closer with respect to the other datasets) are highlighted. Error bars: s.e.m. across the 10 models (each averaged over 5 runs).
Figure 4: Distance to the training data and the signal tracked by trained confidence. Each point is one test question for one training dataset; its y-value is Δe computed from accuracy, consistency and confidence averaged over 10 models × 5 training runs. Negative values mean confidence is closer to true accuracy, positive values closer to output consistency. Points are colored by training dataset (A) or test dataset (B). The x-axis is the mean cosine distance to the 100 nearest training questions. Black line: linear fit ( r=0.49 ); dashed: Δe=0 ; 0.13% points beyond ±0.4 not shown.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Level
Operators
b range
a range
#questions per operator
easy
+,−,×
[1,9]
[10,99999]
12000
medium
+,−,×
[10,99]
[100,99999]
16000
hard
+,−,×
[100,999]
[1000,99999]
12000
Appendix
Table 1: Categories of the Math datasets. Each level is split evenly between the three operators, yielding 9 categories (e.g., easy + , easy − , easy ∗ ).
easy − , easy + , medium − , medium ∗ , medium + , hard − , hard ∗ , hard +
test (1)
easy ∗
Appendix
Table 2: Category-level train/test split of each dataset, selected by the procedure of Appendix C . The 4- and 10-option variants of MMLU-Pro and Math share the same split. Counts in parentheses are the number of categories per split.
Model
MMLU-Pro 4
MMLU-Pro 10
MedMCQA 4
Math 4
Math 10
train
test
train
test
train
test
train
test
train
test
#items ( ND )
10 685 / 1 256
10 337 / 1 202
108 642 / 12 123
108 000 / 12 000
108 000 / 12 000
Llama-2-7b-chat
35.5
33.1
18.3
17.3
37.8
36.0
31.8
32.2
16.3
19.2
Llama-3.2-3B-Instruct
42.1
41.8
25.6
25.5
66.9
65.6
34.1
28.6
13.7
12.8
Mistral-7B-Instruct-v0.3
46.1
48.1
28.3
29.0
52.1
50.6
31.2
26.9
12.8
13.5
Ministral-3-3B-Instruct
46.3
47.2
31.1
31.9
49.7
49.8
37.1
32.5
15.8
16.1
Appendix
Table 3: Accuracy (%) of the evaluated models on the train and test splits of each dataset variant. Mean averages over the 10 models, and ∣Δ∣ is the absolute difference between the train and test means. This is a looser measure than the per-model gap minimized by the splitting objective (Appendix C ). #items is given as train / test. MMLU-Pro 10 contains fewer items than MMLU-Pro 4 because of the 500-token length limit: with 10 options, some questions exceed the limit that they satisfy with 4 options.
Hyperparameter
Value
Optimizer
Paged AdamW 8bits
Learning rate
5×10−5
LR schedule
Linear
Warmup ratio
0.1
Weight decay
10−2
Batch size
16
Appendix
Table 4: Hyperparameters used for calibration training. The same values are used for all models and datasets.
Figure 5: Development of output consistency tracking during training. Each column shows one training dataset. The top row represents Δr over training; positive values indicate that confidence tracks true accuracy more than output consistency, negative values the reverse. Bottom row represents Pearson correlation between trained confidence and output consistency. Curves are computed on the train and test splits of the training dataset. Shaded areas: s.e.m. across the 10 models. For this specific experiment models were trained on MedMCQA for 8 epochs instead of 1 to investigate the overfitting regime.
Figure 6: Relationship between predicted accuracy, true accuracy and output consistency with middle probe. Left panel shows the trained confidence of Qwen3.5-9B against true accuracy (top) and output consistency (bottom) on the test split of MedMCQA 4 and Math 4 with the probe located at the middle of the network. Orange lines are quadratic fits ( R2 shown above each plot). Right panel represents Δr on each test split for both middle and end probes; positive values (red) indicate that confidence tracks true accuracy more than output consistency, negative values (blue) the reverse. Error bars: s.e.m. across the 10 models (each averaged over 5 runs).
Figure 7: Cross-dataset evaluation of trained confidence with the middle probe. Each panel shows models trained on one dataset and evaluated on the test split of all five datasets. Bars represent Δr ; positive values (red) indicate that confidence tracks true accuracy more than output consistency, negative values (blue) the reverse. In-distribution evaluations (tested on the test split from the training dataset - thematically distinct but closer with respect to the other datasets) are highlighted. Error bars: s.e.m. across the 10 models (each averaged over 5 runs).
Figure 8: Development of output consistency tracking during training with the middle probe. Each column shows one training dataset. The top row represents Δr over training; positive values indicate that confidence tracks true accuracy more than output consistency, negative values the reverse. Bottom row represents Pearson correlation between trained confidence and output consistency. Curves are computed on the train and test splits of the training dataset. Shaded areas: s.e.m. across the 10 models.
Figure 9: Distance to the training data and the signal tracked by trained confidence with the mid probe. Each point is one test question for one training dataset; its y-value is Δe computed from accuracy, consistency and confidence averaged over 10 models × 5 training runs. Negative values mean confidence is closer to true accuracy, positive values closer to output consistency. Points are colored by training dataset (A) or test dataset (B). The x-axis is the mean cosine distance to the 100 nearest training questions. Black line: linear fit ( r=0.07 ); dashed: Δe=0 ; 0.10% of points beyond ±0.4 not shown.
Figure 10: Loss during training for both end and mid probes, train and test sets.
Figure 11: Confidence performances before and after training on test splits for each dataset. Bars represent the pearson r between accuracy and verbalised predicted accuracy (blue) or trained predicted accuracy (red) on each of the 5 datasets and the 10 models. Each bar is the average of 5 runs on the same dataset x model and errorbars represent the stderr between these trainings.
Figure 12: Confidence performances before and after training on test splits for each dataset on both probes. Bars represent the pearson r between accuracy and verbalised predicted accuracy (blue), trained predicted accuracy on the end probe (red) or trained predicted accuracy on the mid probe (orange) on each of the 5 datasets and the 10 models. Each bar is the average of 5 runs on the same dataset x model and errorbars represent the stderr between these trainings.
Figure 13: Trained confidence vs. accuracy for all models and datasets with the end probe. X axis is the datasets, Y axis the model. Line is a 2nd order polynomial fitted on the data, R2 its fit coefficient and r is the Pearson correlation coefficient.
Figure 14: Trained confidence vs. consistency for all models and datasets with the end probe. X axis is the datasets, Y axis the model. Line is a 2nd order polynomial fitted on the data, R2 its fit coefficient and r is the Pearson correlation coefficient.
Figure 15: Trained confidence vs. accuracy for all models and datasets with the mid probe. X axis is the datasets, Y axis the model. Line is a 2nd order polynomial fitted on the data, R2 its fit coefficient and r is the Pearson correlation coefficient.
Figure 16: Trained confidence vs. consistency for all models and datasets with the mid probe. X axis is the datasets, Y axis the model. Line is a 2nd order polynomial fitted on the data, R2 its fit coefficient and r is the Pearson correlation coefficient.
Figure 17: Cross dataset true accuracy correlation with trained confidence across both probes. Left column represents end probe results, right column, mid probe results and bottom plots show average + s.e.m. on all models together.
Figure 18: Cross dataset output consistency correlation with trained confidence across both probes. Left column represents end probe results, right column, mid probe results and bottom plots show average + s.e.m. on all models together.
Figure 19: Cross dataset Δr across both probes. Left column represents end probe results, right column, mid probe results and bottom plots show average + s.e.m. on all models together.
Confidence-weighted routing, selective abstention, and ensemble weighting all assume that a model's stated confidence is informative about its capability on the question being asked. They presume functional metacognition, the capacity to assess one's own capabilities, without exercising them. Aggregate calibration is well studied, with mixed results, but the underlying structure of elicited confidence is less well understood. We decompose binary confidence judgements from 20 frontier Large Language Models (LLMs) across six benchmarks using tetrachoric factor analysis paired with pairwise calibration, asking whether two models that differ in confidence also differ in performance. On factual recall and information retrieval benchmarks the cross-model confidence matrix is approximately rank-one and a single dominant factor captures most of the latent variance. Models retrieving facts share an item-level difficulty axis and differ mainly in their decision thresholds along it. Across all benchmarks the relationship between confidence and performance collapses once items that all models agree on are removed. Inter-model pairwise calibration is small even where statistically significant, and what remains shrinks to nothing once base-rate differences along the shared factor are controlled for. Mathematical reasoning is the apparent exception, but this turns out to be a confound where reasoning models answer questions about their confidence by trying to solve them in their chain of thought, bypassing the sub-symbolic self-knowledge we seek to measure. We find no evidence for significant verbalised individuated metacognition in any tested domain.
Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own cognitive processes. Yet LLMs exhibit systemic deficiencies in key metacognitive faculties: they hallucinate with high confidence, fail to recognize knowledge boundaries, and misrepresent their internal uncertainty--undermining trustworthiness and reliability. Since monitoring task performance and adapting behavior accordingly are central to metacognition, we posit that models capable of accurately judging their own performance are better positioned to improve it. We operationalize this idea via two novel mechanisms: reinforcement learning with metacognitive feedback (RLMF), a paradigm to refine completion rankings during preference optimization based on the quality of a model's self-judgments of performance, and metacognitive data selection, which uses similar self-judgments to identify high-value training examples, outperforming naive active learning. We apply these innovations to the problem of faithful calibration (FC), a task that is itself fundamentally metacognitive: the goal is to align expressed with intrinsic uncertainty, difficult even for frontier LLMs. We adopt a two-stage, decoupled approach, first using these methods to calibrate the faithfulness of models' self-reported confidence scores, then mapping to natural, context-adaptable linguistic uncertainty via targeted output editing. Extensive experiments show RLMF achieves generalizable, state-of-the-art FC on diverse tasks while preserving accuracy. Further, RLMF surpasses standard RL by up to 63% while enhancing models' ability to assess and express their own capability limits. This positions RLMF as a promising paradigm to enhance LLM metacognition toward improved abilities and alignment, and suggests metacognitive performance as an effective RL signal to overcome limits of prior intrinsic feedback methods.
Gabrielle Kaili-May Liu, Avi Caciularu, Gal Yona +2
Reliable decisions depend on recognizing when an answer may be wrong. In biological cognition, metacognitive monitoring can dissociate from task performance, raising the question of how closely solving and judging are linked in language models. Here we study the confidence reports of four frontier models across 15 benchmarks. High task accuracy can coexist with weak error discrimination: a model solves 97% of competition mathematics problems while its answer-time confidence ranks correct answers above errors barely better than chance. Confidence separates correct answers from errors more effectively on questions solved by a separate reference model, while review brings limited improvement on reference-hard questions. Aggregate discrimination also rewards ranking correct answers on easy questions above errors on hard ones, which question-only forecasts already do well. Cross-evaluation helps most where the evaluator answered correctly, and errors shared by the two models usually retain high confidence. Hard questions and shared errors remain difficult targets for prompted self-review and peer oversight, even in models with strong problem-solving performance.