Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, and that fine-tuning can substantially improve them. However, what models actually learn during such training remains poorly understood. We investigate how LLMs acquire metacognitive monitoring, the ability to know what one knows, by training 10 open-weight LLMs to predict their own accuracy on factual multiple-choice questions before answering them. We find that trained confidence reflects two distinct signals. While on questions close to the training data, it tracks the model's true accuracy, in other domains, it instead tracks output consistency: the concentration of the model's answer distribution. Output consistency tracking emerges early in training and generalizes across datasets, whereas accuracy tracking develops later and remains local to the training distribution. These results suggest that calibration training may not teach models to generally detect errors they commit confidently, and they raise broader questions about the nature of metacognition in artificial systems.
Figures & tables
Figure 1: Trained confidence tracks accuracy locally and consistency globally Left side: experimental setup — models answer factual multiple-choice questions from five dataset variants. Accuracy is the probability assigned to the correct option, and output consistency is the normalized concentration of the answer distribution; the two can differ sharply, e.g., when a model reliably selects a wrong answer. Before training, verbalized confidence is essentially uncorrelated with true accuracy ( r=0.04 ). After fine-tuning with LoRA and a linear probe, trained confidence predicts accuracy much better ( r=0.73 ). Right side: schematics of the results — trained confidence tracks true accuracy on questions near the training distribution but output consistency on more distant ones, such as held-out topics or other datasets.
Figure 2: Relationship between predicted accuracy, true accuracy and output consistency. Left panel shows the trained confidence of Qwen3.5-9B against true accuracy (top) and output consistency (bottom) on the test split of MedMCQA 4 and Math 4 . Red lines are quadratic fits ( R2 shown above each plot). Right panel represents Δr on each test split; positive values (red) indicate that confidence tracks true accuracy more than output consistency, negative values (blue) the reverse. Error bars: s.e.m. across the 10 models (each averaged over 5 runs).
Figure 3: Cross-dataset evaluation of trained confidence. Each panel shows models trained on one dataset and evaluated on the test split of all five datasets. Bars represent Δr ; positive values (red) indicate that confidence tracks true accuracy more than output consistency, negative values (blue) the reverse. In-distribution evaluations (tested on the test split from the training dataset - thematically distinct but closer with respect to the other datasets) are highlighted. Error bars: s.e.m. across the 10 models (each averaged over 5 runs).
Figure 4: Distance to the training data and the signal tracked by trained confidence. Each point is one test question for one training dataset; its y-value is Δe computed from accuracy, consistency and confidence averaged over 10 models × 5 training runs. Negative values mean confidence is closer to true accuracy, positive values closer to output consistency. Points are colored by training dataset (A) or test dataset (B). The x-axis is the mean cosine distance to the 100 nearest training questions. Black line: linear fit ( r=0.49 ); dashed: Δe=0 ; 0.13% points beyond ±0.4 not shown.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Level
Operators
b range
a range
#questions per operator
easy
+,−,×
[1,9]
[10,99999]
12000
medium
+,−,×
[10,99]
[100,99999]
16000
hard
+,−,×
[100,999]
[1000,99999]
12000
Appendix
Table 1: Categories of the Math datasets. Each level is split evenly between the three operators, yielding 9 categories (e.g., easy + , easy − , easy ∗ ).
easy − , easy + , medium − , medium ∗ , medium + , hard − , hard ∗ , hard +
test (1)
easy ∗
Appendix
Table 2: Category-level train/test split of each dataset, selected by the procedure of Appendix C . The 4- and 10-option variants of MMLU-Pro and Math share the same split. Counts in parentheses are the number of categories per split.
Model
MMLU-Pro 4
MMLU-Pro 10
MedMCQA 4
Math 4
Math 10
train
test
train
test
train
test
train
test
train
test
#items ( ND )
10 685 / 1 256
10 337 / 1 202
108 642 / 12 123
108 000 / 12 000
108 000 / 12 000
Llama-2-7b-chat
35.5
33.1
18.3
17.3
37.8
36.0
31.8
32.2
16.3
19.2
Llama-3.2-3B-Instruct
42.1
41.8
25.6
25.5
66.9
65.6
34.1
28.6
13.7
12.8
Mistral-7B-Instruct-v0.3
46.1
48.1
28.3
29.0
52.1
50.6
31.2
26.9
12.8
13.5
Ministral-3-3B-Instruct
46.3
47.2
31.1
31.9
49.7
49.8
37.1
32.5
15.8
16.1
Appendix
Table 3: Accuracy (%) of the evaluated models on the train and test splits of each dataset variant. Mean averages over the 10 models, and ∣Δ∣ is the absolute difference between the train and test means. This is a looser measure than the per-model gap minimized by the splitting objective (Appendix C ). #items is given as train / test. MMLU-Pro 10 contains fewer items than MMLU-Pro 4 because of the 500-token length limit: with 10 options, some questions exceed the limit that they satisfy with 4 options.
Hyperparameter
Value
Optimizer
Paged AdamW 8bits
Learning rate
5×10−5
LR schedule
Linear
Warmup ratio
0.1
Weight decay
10−2
Batch size
16
Appendix
Table 4: Hyperparameters used for calibration training. The same values are used for all models and datasets.
Figure 5: Development of output consistency tracking during training. Each column shows one training dataset. The top row represents Δr over training; positive values indicate that confidence tracks true accuracy more than output consistency, negative values the reverse. Bottom row represents Pearson correlation between trained confidence and output consistency. Curves are computed on the train and test splits of the training dataset. Shaded areas: s.e.m. across the 10 models. For this specific experiment models were trained on MedMCQA for 8 epochs instead of 1 to investigate the overfitting regime.
Figure 6: Relationship between predicted accuracy, true accuracy and output consistency with middle probe. Left panel shows the trained confidence of Qwen3.5-9B against true accuracy (top) and output consistency (bottom) on the test split of MedMCQA 4 and Math 4 with the probe located at the middle of the network. Orange lines are quadratic fits ( R2 shown above each plot). Right panel represents Δr on each test split for both middle and end probes; positive values (red) indicate that confidence tracks true accuracy more than output consistency, negative values (blue) the reverse. Error bars: s.e.m. across the 10 models (each averaged over 5 runs).
Figure 7: Cross-dataset evaluation of trained confidence with the middle probe. Each panel shows models trained on one dataset and evaluated on the test split of all five datasets. Bars represent Δr ; positive values (red) indicate that confidence tracks true accuracy more than output consistency, negative values (blue) the reverse. In-distribution evaluations (tested on the test split from the training dataset - thematically distinct but closer with respect to the other datasets) are highlighted. Error bars: s.e.m. across the 10 models (each averaged over 5 runs).
Figure 8: Development of output consistency tracking during training with the middle probe. Each column shows one training dataset. The top row represents Δr over training; positive values indicate that confidence tracks true accuracy more than output consistency, negative values the reverse. Bottom row represents Pearson correlation between trained confidence and output consistency. Curves are computed on the train and test splits of the training dataset. Shaded areas: s.e.m. across the 10 models.
Figure 9: Distance to the training data and the signal tracked by trained confidence with the mid probe. Each point is one test question for one training dataset; its y-value is Δe computed from accuracy, consistency and confidence averaged over 10 models × 5 training runs. Negative values mean confidence is closer to true accuracy, positive values closer to output consistency. Points are colored by training dataset (A) or test dataset (B). The x-axis is the mean cosine distance to the 100 nearest training questions. Black line: linear fit ( r=0.07 ); dashed: Δe=0 ; 0.10% of points beyond ±0.4 not shown.
Figure 10: Loss during training for both end and mid probes, train and test sets.
Figure 11: Confidence performances before and after training on test splits for each dataset. Bars represent the pearson r between accuracy and verbalised predicted accuracy (blue) or trained predicted accuracy (red) on each of the 5 datasets and the 10 models. Each bar is the average of 5 runs on the same dataset x model and errorbars represent the stderr between these trainings.
Figure 12: Confidence performances before and after training on test splits for each dataset on both probes. Bars represent the pearson r between accuracy and verbalised predicted accuracy (blue), trained predicted accuracy on the end probe (red) or trained predicted accuracy on the mid probe (orange) on each of the 5 datasets and the 10 models. Each bar is the average of 5 runs on the same dataset x model and errorbars represent the stderr between these trainings.
Figure 13: Trained confidence vs. accuracy for all models and datasets with the end probe. X axis is the datasets, Y axis the model. Line is a 2nd order polynomial fitted on the data, R2 its fit coefficient and r is the Pearson correlation coefficient.
Figure 14: Trained confidence vs. consistency for all models and datasets with the end probe. X axis is the datasets, Y axis the model. Line is a 2nd order polynomial fitted on the data, R2 its fit coefficient and r is the Pearson correlation coefficient.
Figure 15: Trained confidence vs. accuracy for all models and datasets with the mid probe. X axis is the datasets, Y axis the model. Line is a 2nd order polynomial fitted on the data, R2 its fit coefficient and r is the Pearson correlation coefficient.
Figure 16: Trained confidence vs. consistency for all models and datasets with the mid probe. X axis is the datasets, Y axis the model. Line is a 2nd order polynomial fitted on the data, R2 its fit coefficient and r is the Pearson correlation coefficient.
Figure 17: Cross dataset true accuracy correlation with trained confidence across both probes. Left column represents end probe results, right column, mid probe results and bottom plots show average + s.e.m. on all models together.
Figure 18: Cross dataset output consistency correlation with trained confidence across both probes. Left column represents end probe results, right column, mid probe results and bottom plots show average + s.e.m. on all models together.
Figure 19: Cross dataset Δr across both probes. Left column represents end probe results, right column, mid probe results and bottom plots show average + s.e.m. on all models together.