cs.CLSep 27, 2026

LLMs learn different forms of metacognition when trained to predict their own accuracy

Authors: Nicolas Yax, Stefano Palminteri, Pierre-Yves Oudeyer

Organizations: LNC2, INSERM, Paris, France · DEC, ENS, PSL, Paris, France · Flowers AI & CogSci Lab · Centre Inria de l’Université de Bordeaux France

Abstract

Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, and that fine-tuning can substantially improve them. However, what models actually learn during such training remains poorly understood. We investigate how LLMs acquire metacognitive monitoring, the ability to know what one knows, by training 10 open-weight LLMs to predict their own accuracy on factual multiple-choice questions before answering them. We find that trained confidence reflects two distinct signals. While on questions close to the training data, it tracks the model's true accuracy, in other domains, it instead tracks output consistency: the concentration of the model's answer distribution. Output consistency tracking emerges early in training and generalizes across datasets, whereas accuracy tracking develops later and remains local to the training distribution. These results suggest that calibration training may not teach models to generally detect errors they commit confidently, and they raise broader questions about the nature of metacognition in artificial systems.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

    Jun 30, 2026Gabrielle Kaili-May Liu, Avi Caciularu, Gal Yona +2MetacognitionLLM Reasoning Strategies

  2. On the Limits of Metacognitive Monitoring in LLMs

    Sep 28, 2026Dongqi Han, Yifan Yang, Dongsheng LiMetacognitionConfidence Calibration