Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
Authors: Yuxi Xia, Dennis Ulmer, Terra Blevins, Yihong Liu, Hinrich Schütze, Benjamin Roth
Organizations: Faculty of Computer Science, UniVie Doctoral School Computer Science · Faculty of Philological and Cultural Studies, University of Vienna, Austria · ILLC, University of Amsterdam, Netherlands · Khoury College of Computer Sciences, Northeastern University, USA · LMU Munich, Munich Center for Machine Learning (MCML), Germany
Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing evaluations mainly concern the alignment between confidence and correctness, but ignore the variability of language: confidence estimates should remain consistent under semantically equivalent prompts or answer variations, while changing when answer meaning differs, as this may indicate a change in correctness. Therefore, we introduce a novel evaluation framework based on three complementary properties: \textbf{robustness} to prompt perturbations, \textbf{stability} across semantically equivalent answers, and \textbf{sensitivity} to semantically different answers. We show that these metrics are largely independent from existing CE metrics, and that common CE methods often fail on them: while most methods achieve high robustness and stability, they struggle to distinguish semantically different answers, potentially because they do not effectively leverage generation-side information. Overall, our framework exposes overlooked limitations of current CE evaluations and provides guidance for selecting confidence estimators for real-world applications.
Figures & tables
Figure 1: Example of different failure cases of confidence estimators under language variation in either the input prompt or the generated response of the LLM.
ECE ↓
Brier ↓
AUROC ↑
A-STB ↑
A-SST ↑
foracle
0.00
0.00
1.00
1.00
1.00
fconstant
0.01
0.17
0.83
1.00
0.00
frandom
0.33
0.33
0.67
0.72
0.18
fprior
0.00
0.25
0.50
1.00
0.00
Table 1: Simulation results. Two of our own metrics are stability (A-STB) and sensitivity (A-SST).
Figure 2: Design of the proposed evaluation metrics: Robustness measures the score variation under semantically irrelevant prompt perturbations. Stability measures confidence consistency across equivalent answers. Sensitivity evaluates whether CE methods assign distinct scores to answers with different meanings than to equivalent ones.
Figure 3: Comparison of the information CE evaluation metrics require. While others need correctness labels, P-RB, A-STB and A-SST solely rely on multiple answers under different prompts or stochastic decoding.
Figure 4: Similarity of CE evaluation metrics. (a) Heatmap of Kendall’s τ between metrics, including accuracy. (b) Hierarchical clustering based on τ values.
Figure 5: Normalized metric value distributions by model family and size. Whiskers and shaded regions indicate the standard deviation over observed metric values across datasets and CE methods. Models of the same family are connected and appear in the same color. Trend lines are second-degree polynomials fitted on all observations.
SciQ
PopQA
Method
ECE ↓
Brier ↓
AUROC ↑
P-RB ↑
A-STB ↑
A-SST ↑
ECE ↓
Brier ↓
AUROC ↑
P-RB ↑
A-STB ↑
A-SST ↑
Logit-based
Seq. Likelihood
0.077
0.150
0.683
0.938
0.919
0.167
0.386
0.309
0.841
0.951
0.936
0.196
Boosted Prob.
0.176
0.180
0.725
0.994
0.996
0.011
0.575
0.510
0.862
0.979
0.973
0.072
Platt Scaling
0.096
0.156
0.683
0.987
0.981
0.039
0.277
0.253
0.841
0.990
0.987
0.040
P(True)
0.153
0.167
0.729
0.953
0.970
0.183
0.329
0.316
0.797
0.928
0.938
0.251
Internal States
Attention Score
0.017
0.152
0.552
0.991
0.993
0.019
0.051
0.199
0.614
0.987
0.981
0.046
Table 2: Calibration and our metrics averaged across all models. Arrows indicate metric direction ( ↓ lower is better, ↑ higher is better); the best value per column is bold. Results of other datasets are shown in Table 3 in Section B.1 .
Figure 6: Mean reciprocal rank of CE methods per evaluation metric, averaged over datasets. ∅ is another average across all metrics, including standard deviation.
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
NQ
TriviaQA
Method
ECE ↓
Brier ↓
AUROC ↑
P-RB ↑
A-STB ↑
A-SST ↑
ECE ↓
Brier ↓
AUROC ↑
P-RB ↑
A-STB ↑
A-SST ↑
Logit-based
Seq. Likelihood
0.404
0.374
0.716
0.952
0.942
0.163
0.275
0.282
0.702
0.950
0.924
0.200
Boosted Prob.
0.568
0.543
0.713
0.988
0.988
0.033
0.383
0.371
0.713
0.992
0.991
0.033
Platt Scaling
0.227
0.273
0.716
0.990
0.988
0.034
0.099
0.236
0.702
0.989
0.983
0.045
P(True)
0.345
0.354
0.722
0.929
0.947
0.201
0.268
0.283
0.700
0.947
0.955
0.278
Internal States
Attention Score
0.044
0.226
0.583
0.987
0.987
0.029
0.033
0.234
0.558
0.988
0.987
0.034
Appendix
Table 3: Calibration and our metrics averaged across all models. Arrows indicate metric direction ( ↓ lower is better, ↑ higher is better); the best value per column is bold.
Comparison
Agreement
Cohen’s κ
Human agreement
Human 1 vs. Human 2
0.930
0.839
Model–human agreement
GPT-4o vs. Human 1
0.930
0.842
GPT-4o vs. Human 2
0.900
0.774
GPT-5.5 vs. Human 1
0.930
0.832
Appendix
Table 4: Agreement on response-correctness judgments among human and LLM evaluators. Human consensus is defined over items for which the two human annotators agreed. All-judge agreement is the proportion of items on which all four evaluators assigned the same label. Cohen’s κ is reported only for pairwise comparisons over the full sample.
Validator
Approval Rate
Human 1
0.975
Human 2
0.965
GPT-5.5
0.905
All validators
0.884
Appendix
Table 5: Validator approval rates for the generated answer partitions.
Figure 7: Accuracy of models on different datasets.
Figure 8: Distribution of metric across models and CE methods, shown per-dataset. Whiskers show the interquartile range and dots signify outliers.
Metric
ECE
Brier
AUROC
P-RB
A-STB
A-SST
Acc.
ECE
1
Brier
0.644∗∗∗
1
[0.590, 0.696]
AUROC
0.303∗∗∗
0.084∗
1
[0.234, 0.368]
[0.019, 0.157]
P-RB
−0.102∗∗
0.016
−0.330∗∗∗
1
Appendix
Table 6: Mean dataset-wise Kendall’s τ correlations among evaluation metrics and model accuracy. Brackets report cluster-bootstrap 95% confidence intervals. Asterisks indicate BH-corrected permutation significance: ∗p<0.05 , ∗∗p<0.01 , and ∗∗∗p<0.001 .
NQ
PopQA
SciQ
TriviaQA
∣I∣/M
0.698
0.643
0.860
0.877
Appendix
Table 7: Average fraction of prompts retained ∣I∣/M after semantic-equivalence filtering.
Method
A-SSTsame
A-SSTdifferent
Δ
Seq. Likelihood
0.183
0.200
0.017
Boosted Prob.
0.029
0.014
−0.015
Platt Scaling
0.039
0.043
0.004
P(True)
0.231
0.347
0.116
Attention Score
0.028
0.033
0.005
Hidden Score
0.040
0.046
0.006
Appendix
Table 8: Mean A-SST for pairs of semantic groups with the same A-SSTsame or different A-SSTdifferent answer-correctness labels. The difference is calculated as Δ=A-SSTdifferent−A-SSTsame .
Figure 9: Normalized metric value distributions by model family and size. Whiskers and shaded regions indicate the standard deviation over observed metric values across datasets and CE methods. Models of the same family are connected and appear in the same color. Trend lines are second degree polynomials fitted on all observations.
Figure 10: Normalized metric value distributions by model family and size. Confidence estimators from the same estimator family and for the same model family are connected and appear in the same color. Measurements from confidence estimators in the same family are shown with the same marker. Trend lines are second-degree polynomials fitted on all observations.
Figure 11: Reciprocal rank of CE methods per evaluation metric over datasets.
The electrode at which oxidation occurs is called? y : the anode. Corr.: 1.00
CE
t1
t2
t3
t4
t5
t6
t7
t8
t9
t10
P-RB
Qwen-14B
SciQ
Seq. Likelihood
0.96
0.72
0.98
1.00
1.00
1.00
0.99
0.89
1.00
1.00
0.92
Boosted Prob.
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
Platt Scaling
0.75
0.75
0.75
0.75
0.75
0.75
0.75
0.73
0.75
0.75
0.99
P(True)
1.00
1.00
0.00
0.00
1.00
1.00
0.00
1.00
0.00
0.00
0.50
Attention Score
0.81
0.81
0.80
0.81
0.81
0.81
0.80
0.80
0.81
0.81
1.00
Appendix
Table 9: Case study 1 of prompt robustness. The confidence estimated by different CE methods under 10 perturbed answer-elicit prompts ( t1 – t10 ) for the same question and LLM response is shown.
With an atomic weight of 22, what element, named for members of Greek mythology, uses the symbol Ti?
y : Titanium ore. Corr.: 1.00
CE
t1
t2
t3
t4
t5
t6
t7
t8
t9
t10
P-RB
Mistral-123B
TriviaQA
Seq. Likelihood
0.99
0.87
0.90
0.92
0.79
1.00
0.91
0.99
0.90
0.84
0.93
Boosted Prob.
1.00
0.99
1.00
0.99
0.97
1.00
1.00
1.00
0.99
0.99
0.99
Platt Scaling
0.73
0.70
0.71
0.71
0.68
0.73
0.71
0.73
0.71
0.70
0.99
P(True)
0.08
0.01
0.05
0.01
0.65
0.08
0.03
0.08
0.02
0.84
0.72
Appendix
Table 10: Case study 2 of prompt robustness. The confidence estimated by different CE methods under 10 perturbed prompts t1 – t10 ) for the same question and LLM response is shown.
What type of lens is thicker at the edges than it is in the middle? y∗ : concave lens
Convex
Concave lens
Concave
CE
Y^1
Y^2
Y^2
A-STB
A-SST
Llama-70B
SciQ
Seq. Likelihood
0.59
0.59
0.82
0.90
0.17
Boosted Prob.
1.00
1.00
1.00
1.00
0.00
Platt Scaling
0.67
0.67
0.73
0.98
0.04
P(True)
0.32
0.00
0.00
1.00
0.32
Appendix
Table 11: Case study 1 of answer stability and sensitivity. The confidence estimated by different CE methods for unique sampled answers to the same question is shown. Purple marks the answer group used for A-STB; purple and green mark the semantic groups used for A-SST.
where was the world chess tournament 2017 held? y∗ : Tbilisi, Georgia
New York City, USA
Riyadh, Saudi Arabia
New York City
Carlsen vs. Caruana 2018
CE
Y^1
Y^2
Y^3
Y^4
A-STB
A-SST
Mistral-123B
NQ
Seq. Likelihood
0.61
0.69
0.74
0.64
0.94
0.03
Boosted Prob.
0.99
0.99
0.99
0.98
1.00
0.00
Platt Scaling
0.60
0.62
0.63
0.61
0.99
0.01
P(True)
0.05
0.87
0.04
0.01
0.99
0.03
Appendix
Table 12: Case study 2 of answer stability and sensitivity. Purple marks the answer group used for A-STB; purple and green mark the semantic groups used for A-SST.
What is the capital of Cao Wei? y∗ : Luoyang
洛阳
Luo Yin
Luoyang
CE
y1
y2
y3
A-STB
A-SST
Qwen-14B
PopQA
Seq. Likelihood
0.47
0.46
0.61
0.94
0.01
Boosted Prob.
1.00
0.97
1.00
1.00
0.03
Platt Scaling
0.49
0.49
0.52
0.99
0.00
P(True)
1.00
0.00
1.00
1.00
1.00
Appendix
Table 13: Case study 3 of answer stability and sensitivity. Purple marks the answer group used for A-STB; purple and green mark the semantic groups used for A-SST. 洛阳 is the Chinese character of “Luoyang”.
Method
ECE ↓
Brier ↓
AUROC ↑
P-RB ↑
A-STB ↑
A-SST ↑
Seq. Likelihood
0.264
0.275
0.678
0.949
0.916
0.216
Platt Scaling
0.076
0.216
0.649
0.990
0.982
0.047
Verbalized Conf.
0.258
0.262
0.701
0.962
0.985
0.109
P(True)
0.295
0.296
0.733
0.976
0.988
0.226
Calib1
0.104
0.208
0.704
0.982
0.960
0.091
Appendix
Table 14: GPT-4o results aggregated over all datasets. Due to the limited access of the commercial model, only methods that do not require information beyond the top 20 output logits apply to this model.
Concept
Prompt Name
Prompt
Scale Variants
P(1)
Provide the probability that your answer is correct. Give ONLY the probability between 0.0 and 1.0 , no other words or explanation.
P(%)
Provide the probability that your answer is correct. Give ONLY the probability between 0% and 100% , no other words or explanation.
P(10)
Provide the probability that your answer is correct. Give ONLY the probability between 0 and 10 , no other words or explanation.
Lexical Variants
CF(1)
Provide the confidence that your answer is correct. Give ONLY the confidence between 0.0 and 1.0, no other words or explanation.
CT(1)
Provide the certainty that your answer is correct. Give ONLY the confidence between 0 and 10, no other words or explanation.
Linguistic Expressions
L.
Describe how likely it is that your answer is correct as one of the following expressions : [’Almost No Chance’, ’Highly Unlikely’, ’Chances are Slight’, ’Little Chance’, ’Unlikely’, ’Probably Not’, ’About Even’, ’Better than Even’, ’Likely’, ’Probably’, ’Very Good Chance’,’Highly Likely’, ’Almost Certain’]. Give ONLY the chosen expression, no other words or explanation.
Appendix
Table 15: Detail of prompts for eliciting confidence from LLMs. These prompts are specifically applied to verbalized confidence, we use these prompts to further validate the robustness of verbalized confidence in Section B.12 .
Dataset
Ans-elicit
Conf-elicit
NQ
0.985
0.907
PopQA
0.984
0.895
SciQ
0.990
0.928
TriviaQA
0.985
0.921
AVG
0.986
0.913
Appendix
Table 16: Robustness of verbalized confidence under two prompt perturbation settings: perturbing answer-elicit prompt (Ans-elicit) and perturbing confidence-elicit prompt (Conf-elicit).
Figure 12: Robustness : Pearson correlations of verbalized confidence when using different confidence-elicit prompts in Table 15 .
Method
Y^min
Y^∈/max
Seq. Likelihood
0.182
0.171
Boosted Prob.
0.037
0.035
Platt Scaling
0.039
0.037
P(True)
0.228
0.227
Attention Score
0.032
0.032
Hidden Score
0.034
0.034
Appendix
Table 17: Comparison of answer sensitivity settings. The original setting compares the largest semantic group with the smallest semantic group Y^min ; the alternative compares the largest group with all remaining semantically different answers Y^∈/max .
Instruction Model
Reasoning Model
Method
ECE
Brier
AUROC
P-RB
A-STB
A-SST
ECE
Brier
AUROC
P-RB
A-STB
A-SST
Seq. Likelihood
0.410
0.387
0.707
0.952
0.975
0.169
0.378
0.321
0.750
0.949
0.938
0.146
Boosted Prob.
0.556
0.536
0.712
0.989
0.996
0.031
0.615
0.568
0.715
0.982
0.976
0.041
Platt Scaling
0.217
0.275
0.707
0.990
0.995
0.035
0.270
0.266
0.750
0.990
0.987
0.030
P(True)
0.332
0.350
0.716
0.925
0.970
0.190
0.401
0.366
0.742
0.947
0.940
0.205
Attention Score
0.045
0.233
0.575
0.987
0.994
0.027
0.042
0.199
0.614
0.986
0.987
0.030
Appendix
Table 18: NQ metric values averaged across instruction and reasoning models, respectively. Darker color indicates better performance.
Instruction Model
Reasoning Model
Method
ECE
Brier
AUROC
P-RB
A-STB
A-SST
ECE
Brier
AUROC
P-RB
A-STB
A-SST
Seq. Likelihood
0.087
0.149
0.677
0.937
0.969
0.172
0.038
0.151
0.705
0.942
0.941
0.147
Boosted Prob.
0.177
0.179
0.723
0.995
0.999
0.010
0.175
0.184
0.732
0.992
0.996
0.017
Platt Scaling
0.097
0.154
0.677
0.986
0.993
0.040
0.092
0.161
0.705
0.987
0.986
0.035
P(True)
0.163
0.170
0.726
0.949
0.992
0.182
0.111
0.153
0.739
0.971
0.989
0.190
Attention Score
0.017
0.150
0.551
0.991
0.997
0.019
0.017
0.160
0.560
0.991
0.996
0.015
Appendix
Table 19: SciQ metric values averaged across instruction and reasoning models, respectively. Darker color indicates better performance.
Instruction Model
Reasoning Model
Method
ECE
Brier
AUROC
P-RB
A-STB
A-SST
ECE
Brier
AUROC
P-RB
A-STB
A-SST
Seq. Likelihood
0.277
0.285
0.688
0.951
0.985
0.206
0.266
0.272
0.757
0.945
0.960
0.173
Boosted Prob.
0.370
0.361
0.701
0.993
0.999
0.029
0.434
0.410
0.761
0.986
0.994
0.049
Platt Scaling
0.079
0.232
0.688
0.990
0.997
0.046
0.177
0.252
0.757
0.988
0.991
0.039
P(True)
0.263
0.279
0.686
0.944
0.992
0.283
0.288
0.297
0.753
0.960
0.986
0.259
Attention Score
0.031
0.230
0.564
0.989
0.998
0.035
0.040
0.249
0.535
0.987
0.994
0.027
Appendix
Table 20: TriviaQA metric values averaged across instruction and reasoning models, respectively. Darker color indicates better performance.
Instruction Model
Reasoning Model
Method
ECE
Brier
AUROC
P-RB
A-STB
A-SST
ECE
Brier
AUROC
P-RB
A-STB
A-SST
Seq. Likelihood
0.407
0.334
0.832
0.950
0.988
0.206
0.301
0.209
0.877
0.955
0.971
0.157
Boosted Prob.
0.584
0.531
0.857
0.982
0.996
0.068
0.537
0.427
0.882
0.968
0.980
0.091
Platt Scaling
0.264
0.257
0.832
0.990
0.998
0.041
0.328
0.236
0.877
0.991
0.994
0.033
P(True)
0.307
0.307
0.797
0.921
0.990
0.257
0.421
0.354
0.795
0.953
0.980
0.229
Attention Score
0.053
0.205
0.612
0.986
0.997
0.049
0.043
0.174
0.623
0.989
0.992
0.033
Appendix
Table 21: PopQA metric values averaged across instruction and reasoning models, respectively. Darker color indicates better performance.
Concept
Prompt Name
Prompt
Answer
A1
Answer the question, give ONLY the answer, no other words or explanation:
A2
Answer the question, give ONLY the answer without explanation:
A3
Answer the question, give ONLY the answer:
A4
Answer the question as short as possible:
A5
Answer the question with minimal words:
Provide
P1
Provide an answer for the question, give ONLY the answer, no other words or explanation:
Appendix
Table 22: Detail of prompts for eliciting answers from LLMs. These are the prompts used for evaluating robustness in the main paper.
Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated output. However, this formulation is misaligned with many practical settings where the central question is how likely a model is to solve a query overall. We show that this mismatch results from the stochastic nature of modern LLM decoding, under which single-response correctness fails to reflect underlying model capability. To address this issue, we introduce capability calibration, a new evaluation framework for measuring how well query-level confidence aligns with a model's expected accuracy on individual queries. We formally distinguish capability calibration (CC) from response calibration (RC) and show that the two differ both theoretically and empirically. We further show that CC is better suited than RC to applications like pass@k prediction and inference budget allocation. Finally, we evaluate common confidence estimation methods to understand the practical feasibility of CC.
LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence. These signals are sometimes treated as direct readouts of model uncertainty, but their comparison depends on measurement choices that are rarely made explicit. In the main analysis, we hold the verbalized-confidence elicitation fixed: a single prompt template, probability scale, and output format. We then vary the measurement axes that define the verbalized-vs-token comparison: which answer string receives the token-probability score, how that score is read from the answer tokens, and under which conditioning context it is measured. We evaluate this design on four QA benchmarks across three open 7--8B base/Instruct model families, with larger Qwen2.5 variants as same-family robustness checks. The resulting comparison is sensitive to these choices: conditioning context changes the sign or magnitude of the ECE gap across settings, token readout produces smaller but still sign-moving changes, and changing the ECE estimator has little effect. Under the default generated-answer, bare-context protocol, Instruct settings are close to parity rather than showing a large calibration gain for verbalized confidence. In a separate supplied-answer analysis, surface-plausible wrong answers receive nearly the same confidence as supplied gold answers, suggesting that verbalized confidence also reflects answer plausibility and provenance rather than correctness alone. We argue that both confidence signals should be treated as protocol-dependent behavioral measurements, and provide a reporting checklist covering elicitation provenance, scored answer, token-probability readout, and conditioning context.
Calibration measures whether a model's predicted confidence aligns with its empirical accuracy, and is central to the reliable deployment of large language models (LLMs) in high-stakes domains such as medicine and law. While much recent work focuses on improving LLM calibration, the equally important question of how to evaluate it in realistic settings remains underdeveloped. Open-ended question answering (QA), the most common deployment setting for modern LLMs, is where existing evaluation methods fall short: logit-based metrics need restricted output formats and internal probabilities; verbalized confidence is self-reported and often overconfident; and sampling-based methods rely on task-specific extraction rules without a clear finite-sample target. We introduce Sem-ECE (Semantic-Sampling Expected Calibration Error), a calibration evaluation framework for open-ended QA that samples answers from the model, groups them into semantic classes, and uses the resulting frequencies as confidence. We study two estimators within this framework: Sem1-ECE, the same-sample self-consistency score, and Sem2-ECE, a held-out variant that separates answer selection from confidence evaluation. We prove both are asymptotically unbiased, and further show that they agree on easy questions but diverge on hard ones with Sem2 achieving strictly smaller calibration error, so their gap also serves as a diagnostic for question difficulty. Experiments on three open-ended QA benchmarks across five leading commercial LLMs match our theoretical predictions and show that Sem-ECE outperforms verbalized confidence and existing sampling-based methods, while complementing logit-based evaluation when internal probabilities are unavailable.
Zhanliang Wang, Jiancong Xiao, Ruochen Jin +3
University of Pennsylvania, Philadelphia, PA · Dartmouth College, Hanover, NH