cs.CLJan 12, 2026

Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

Authors: Yuxi Xia, Dennis Ulmer, Terra Blevins, Yihong Liu, Hinrich Schütze, Benjamin Roth

Organizations: Faculty of Computer Science, UniVie Doctoral School Computer Science · Faculty of Philological and Cultural Studies, University of Vienna, Austria · ILLC, University of Amsterdam, Netherlands · Khoury College of Computer Sciences, Northeastern University, USA · LMU Munich, Munich Center for Machine Learning (MCML), Germany

Abstract

Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing evaluations mainly concern the alignment between confidence and correctness, but ignore the variability of language: confidence estimates should remain consistent under semantically equivalent prompts or answer variations, while changing when answer meaning differs, as this may indicate a change in correctness. Therefore, we introduce a novel evaluation framework based on three complementary properties: \textbf{robustness} to prompt perturbations, \textbf{stability} across semantically equivalent answers, and \textbf{sensitivity} to semantically different answers. We show that these metrics are largely independent from existing CE metrics, and that common CE methods often fail on them: while most methods achieve high robustness and stability, they struggle to distinguish semantically different answers, potentially because they do not effectively leverage generation-side information. Overall, our framework exposes overlooked limitations of current CE evaluations and provides guidance for selecting confidence estimators for real-world applications.

Figures & tables

Appendix figures & tables28 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. On Calibration of Large Language Models: From Response To Capability

    Feb 14, 2026Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin +3Confidence Estimation

  2. Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration

    May 26, 2026Hankyeol Kim, Pilsung KangQuestionProvenance

  3. A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering

    May 8, 2026Zhanliang Wang, Jiancong Xiao, Ruochen Jin +3Confidence EstimationLarge Language Model Evaluation