cs.LGSep 30, 2026

Also Small Models Can Reasonably Self-Evaluate Their Confidence

Authors: Idil Kapikiran, Thomas Decker, Thomas Runkler

Organizations: Technical University of Munich · Siemens AG · LMU Munich · Munich Center for Machine Learning (MCML)

Abstract

This study systematically evaluates self-evaluation-based uncertainty quantification across different language models of varying sizes on question-answering tasks spanning general to specialized knowledge domains. Using various self-evaluation methods where models judge their own predictions, we examine how model scale and domain specificity affect the quality of self-assessed confidence signals. Our results reveal that while accuracy predictably declines with smaller models and more specialized domains, the reliability of self-evaluated confidence remains largely stable across both dimensions. This independence means the most capable model is not necessarily the best at self-assessing prediction reliability. These findings suggest that smaller models can achieve reasonable self-assessed confidence despite lower accuracy, making them viable for resource-constrained deployments.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs

    May 8, 2026Sree Bhattacharyya, Samarth Khanna, Leona Chen +3Self-EvaluationLarge Language Model Reliability

  2. Latent Confidence Alignment for LLM Self-Assessment

    Jun 20, 2026Ting-Yu Chen, Tingting Yu, Pei-Cing Huang +3Self-EvaluationConfidence Calibration