Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
Organizations: Faculty of Computer Science, UniVie Doctoral School Computer Science · Faculty of Philological and Cultural Studies, University of Vienna, Austria · ILLC, University of Amsterdam, Netherlands · Khoury College of Computer Sciences, Northeastern University, USA · LMU Munich, Munich Center for Machine Learning (MCML), Germany
Abstract
Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing evaluations mainly concern the alignment between confidence and correctness, but ignore the variability of language: confidence estimates should remain consistent under semantically equivalent prompts or answer variations, while changing when answer meaning differs, as this may indicate a change in correctness. Therefore, we introduce a novel evaluation framework based on three complementary properties: \textbf{robustness} to prompt perturbations, \textbf{stability} across semantically equivalent answers, and \textbf{sensitivity} to semantically different answers. We show that these metrics are largely independent from existing CE metrics, and that common CE methods often fail on them: while most methods achieve high robustness and stability, they struggle to distinguish semantically different answers, potentially because they do not effectively leverage generation-side information. Overall, our framework exposes overlooked limitations of current CE evaluations and provides guidance for selecting confidence estimators for real-world applications.
Figures & tables
| ECE | Brier | AUROC | A-STB | A-SST | |
|---|---|---|---|---|---|
| 0.00 | 0.00 | 1.00 | 1.00 | 1.00 | |
| 0.01 | 0.17 | 0.83 | 1.00 | 0.00 | |
| 0.33 | 0.33 | 0.67 | 0.72 | 0.18 | |
| 0.00 | 0.25 | 0.50 | 1.00 | 0.00 |
| SciQ | PopQA | ||||||||||||
| Method | ECE | Brier | AUROC | P-RB | A-STB | A-SST | ECE | Brier | AUROC | P-RB | A-STB | A-SST | |
| Logit-based | Seq. Likelihood | 0.150 | |||||||||||
| Boosted Prob. | 0.994 | 0.996 | |||||||||||
| Platt Scaling | 0.987 | 0.990 | |||||||||||
| P(True) | 0.251 | ||||||||||||
| Internal States | Attention Score | 0.017 | 0.152 | 0.991 | 0.051 | 0.987 | |||||||
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| NQ | TriviaQA | ||||||||||||
| Method | ECE | Brier | AUROC | P-RB | A-STB | A-SST | ECE | Brier | AUROC | P-RB | A-STB | A-SST | |
| Logit-based | Seq. Likelihood | ||||||||||||
| Boosted Prob. | 0.988 | 0.713 | 0.992 | ||||||||||
| Platt Scaling | 0.990 | 0.989 | |||||||||||
| P(True) | 0.201 | 0.278 | |||||||||||
| Internal States | Attention Score | 0.044 | 0.987 | 0.234 | 0.988 | ||||||||
| Comparison | Agreement | Cohen’s |
|---|---|---|
| Human agreement | ||
| Human 1 vs. Human 2 | 0.930 | 0.839 |
| Model–human agreement | ||
| GPT-4o vs. Human 1 | 0.930 | 0.842 |
| GPT-4o vs. Human 2 | 0.900 | 0.774 |
| GPT-5.5 vs. Human 1 | 0.930 | 0.832 |
| Validator | Approval Rate |
|---|---|
| Human 1 | 0.975 |
| Human 2 | 0.965 |
| GPT-5.5 | 0.905 |
| All validators | 0.884 |
| Metric | ECE | Brier | AUROC | P-RB | A-STB | A-SST | Acc. |
|---|---|---|---|---|---|---|---|
| ECE | 1 | ||||||
| Brier | 1 | ||||||
| [0.590, 0.696] | |||||||
| AUROC | 1 | ||||||
| [0.234, 0.368] | [0.019, 0.157] | ||||||
| P-RB | 1 |
| NQ | PopQA | SciQ | TriviaQA | |
| 0.698 | 0.643 | 0.860 | 0.877 |
| Method | |||
|---|---|---|---|
| Seq. Likelihood | 0.183 | 0.200 | 0.017 |
| Boosted Prob. | 0.029 | 0.014 | |
| Platt Scaling | 0.039 | 0.043 | 0.004 |
| P(True) | 0.231 | 0.347 | 0.116 |
| Attention Score | 0.028 | 0.033 | 0.005 |
| Hidden Score | 0.040 | 0.046 | 0.006 |
| The electrode at which oxidation occurs is called? : the anode. Corr.: 1.00 | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CE | P-RB | ||||||||||||
| Qwen-14B | SciQ | Seq. Likelihood | 0.96 | 0.72 | 0.98 | 1.00 | 1.00 | 1.00 | 0.99 | 0.89 | 1.00 | 1.00 | 0.92 |
| Boosted Prob. | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | ||
| Platt Scaling | 0.75 | 0.75 | 0.75 | 0.75 | 0.75 | 0.75 | 0.75 | 0.73 | 0.75 | 0.75 | 0.99 | ||
| P(True) | 1.00 | 1.00 | 0.00 | 0.00 | 1.00 | 1.00 | 0.00 | 1.00 | 0.00 | 0.00 | 0.50 | ||
| Attention Score | 0.81 | 0.81 | 0.80 | 0.81 | 0.81 | 0.81 | 0.80 | 0.80 | 0.81 | 0.81 | 1.00 | ||
| With an atomic weight of 22, what element, named for members of Greek mythology, uses the symbol Ti? | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| : Titanium ore. Corr.: 1.00 | |||||||||||||
| CE | P-RB | ||||||||||||
| Mistral-123B | TriviaQA | Seq. Likelihood | 0.99 | 0.87 | 0.90 | 0.92 | 0.79 | 1.00 | 0.91 | 0.99 | 0.90 | 0.84 | 0.93 |
| Boosted Prob. | 1.00 | 0.99 | 1.00 | 0.99 | 0.97 | 1.00 | 1.00 | 1.00 | 0.99 | 0.99 | 0.99 | ||
| Platt Scaling | 0.73 | 0.70 | 0.71 | 0.71 | 0.68 | 0.73 | 0.71 | 0.73 | 0.71 | 0.70 | 0.99 | ||
| P(True) | 0.08 | 0.01 | 0.05 | 0.01 | 0.65 | 0.08 | 0.03 | 0.08 | 0.02 | 0.84 | 0.72 | ||
| What type of lens is thicker at the edges than it is in the middle? : concave lens | |||||||
|---|---|---|---|---|---|---|---|
| Convex | Concave lens | Concave | |||||
| CE | A-STB | A-SST | |||||
| Llama-70B | SciQ | Seq. Likelihood | 0.59 | 0.59 | 0.82 | 0.90 | 0.17 |
| Boosted Prob. | 1.00 | 1.00 | 1.00 | 1.00 | 0.00 | ||
| Platt Scaling | 0.67 | 0.67 | 0.73 | 0.98 | 0.04 | ||
| P(True) | 0.32 | 0.00 | 0.00 | 1.00 | 0.32 | ||
| where was the world chess tournament 2017 held? : Tbilisi, Georgia | ||||||||
|---|---|---|---|---|---|---|---|---|
| New York City, USA | Riyadh, Saudi Arabia | New York City | Carlsen vs. Caruana 2018 | |||||
| CE | A-STB | A-SST | ||||||
| Mistral-123B | NQ | Seq. Likelihood | 0.61 | 0.69 | 0.74 | 0.64 | 0.94 | 0.03 |
| Boosted Prob. | 0.99 | 0.99 | 0.99 | 0.98 | 1.00 | 0.00 | ||
| Platt Scaling | 0.60 | 0.62 | 0.63 | 0.61 | 0.99 | 0.01 | ||
| P(True) | 0.05 | 0.87 | 0.04 | 0.01 | 0.99 | 0.03 | ||
| What is the capital of Cao Wei? : Luoyang | |||||||
|---|---|---|---|---|---|---|---|
| 洛阳 | Luo Yin | Luoyang | |||||
| CE | A-STB | A-SST | |||||
| Qwen-14B | PopQA | Seq. Likelihood | 0.47 | 0.46 | 0.61 | 0.94 | 0.01 |
| Boosted Prob. | 1.00 | 0.97 | 1.00 | 1.00 | 0.03 | ||
| Platt Scaling | 0.49 | 0.49 | 0.52 | 0.99 | 0.00 | ||
| P(True) | 1.00 | 0.00 | 1.00 | 1.00 | 1.00 | ||
| Method | ECE | Brier | AUROC | P-RB | A-STB | A-SST |
|---|---|---|---|---|---|---|
| Seq. Likelihood | 0.264 | 0.275 | 0.678 | 0.949 | 0.916 | 0.216 |
| Platt Scaling | 0.076 | 0.216 | 0.649 | 0.990 | 0.982 | 0.047 |
| Verbalized Conf. | 0.258 | 0.262 | 0.701 | 0.962 | 0.985 | 0.109 |
| P(True) | 0.295 | 0.296 | 0.733 | 0.976 | 0.988 | 0.226 |
| Calib1 | 0.104 | 0.208 | 0.704 | 0.982 | 0.960 | 0.091 |
| Concept | Prompt Name | Prompt |
|---|---|---|
| Scale Variants | P(1) | Provide the probability that your answer is correct. Give ONLY the probability between 0.0 and 1.0 , no other words or explanation. |
| P(%) | Provide the probability that your answer is correct. Give ONLY the probability between 0% and 100% , no other words or explanation. | |
| P(10) | Provide the probability that your answer is correct. Give ONLY the probability between 0 and 10 , no other words or explanation. | |
| Lexical Variants | CF(1) | Provide the confidence that your answer is correct. Give ONLY the confidence between 0.0 and 1.0, no other words or explanation. |
| CT(1) | Provide the certainty that your answer is correct. Give ONLY the confidence between 0 and 10, no other words or explanation. | |
| Linguistic Expressions | L. | Describe how likely it is that your answer is correct as one of the following expressions : [’Almost No Chance’, ’Highly Unlikely’, ’Chances are Slight’, ’Little Chance’, ’Unlikely’, ’Probably Not’, ’About Even’, ’Better than Even’, ’Likely’, ’Probably’, ’Very Good Chance’,’Highly Likely’, ’Almost Certain’]. Give ONLY the chosen expression, no other words or explanation. |
| Dataset | Ans-elicit | Conf-elicit |
|---|---|---|
| NQ | 0.985 | 0.907 |
| PopQA | 0.984 | 0.895 |
| SciQ | 0.990 | 0.928 |
| TriviaQA | 0.985 | 0.921 |
| AVG | 0.986 | 0.913 |
| Method | ||
|---|---|---|
| Seq. Likelihood | 0.182 | 0.171 |
| Boosted Prob. | 0.037 | 0.035 |
| Platt Scaling | 0.039 | 0.037 |
| P(True) | 0.228 | 0.227 |
| Attention Score | 0.032 | 0.032 |
| Hidden Score | 0.034 | 0.034 |
| Instruction Model | Reasoning Model | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | ECE | Brier | AUROC | P-RB | A-STB | A-SST | ECE | Brier | AUROC | P-RB | A-STB | A-SST |
| Seq. Likelihood | 0.410 | 0.387 | 0.707 | 0.952 | 0.975 | 0.169 | 0.378 | 0.321 | 0.750 | 0.949 | 0.938 | 0.146 |
| Boosted Prob. | 0.556 | 0.536 | 0.712 | 0.989 | 0.996 | 0.031 | 0.615 | 0.568 | 0.715 | 0.982 | 0.976 | 0.041 |
| Platt Scaling | 0.217 | 0.275 | 0.707 | 0.990 | 0.995 | 0.035 | 0.270 | 0.266 | 0.750 | 0.990 | 0.987 | 0.030 |
| P(True) | 0.332 | 0.350 | 0.716 | 0.925 | 0.970 | 0.190 | 0.401 | 0.366 | 0.742 | 0.947 | 0.940 | 0.205 |
| Attention Score | 0.045 | 0.233 | 0.575 | 0.987 | 0.994 | 0.027 | 0.042 | 0.199 | 0.614 | 0.986 | 0.987 | 0.030 |
| Instruction Model | Reasoning Model | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | ECE | Brier | AUROC | P-RB | A-STB | A-SST | ECE | Brier | AUROC | P-RB | A-STB | A-SST |
| Seq. Likelihood | 0.087 | 0.149 | 0.677 | 0.937 | 0.969 | 0.172 | 0.038 | 0.151 | 0.705 | 0.942 | 0.941 | 0.147 |
| Boosted Prob. | 0.177 | 0.179 | 0.723 | 0.995 | 0.999 | 0.010 | 0.175 | 0.184 | 0.732 | 0.992 | 0.996 | 0.017 |
| Platt Scaling | 0.097 | 0.154 | 0.677 | 0.986 | 0.993 | 0.040 | 0.092 | 0.161 | 0.705 | 0.987 | 0.986 | 0.035 |
| P(True) | 0.163 | 0.170 | 0.726 | 0.949 | 0.992 | 0.182 | 0.111 | 0.153 | 0.739 | 0.971 | 0.989 | 0.190 |
| Attention Score | 0.017 | 0.150 | 0.551 | 0.991 | 0.997 | 0.019 | 0.017 | 0.160 | 0.560 | 0.991 | 0.996 | 0.015 |
| Instruction Model | Reasoning Model | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | ECE | Brier | AUROC | P-RB | A-STB | A-SST | ECE | Brier | AUROC | P-RB | A-STB | A-SST |
| Seq. Likelihood | 0.277 | 0.285 | 0.688 | 0.951 | 0.985 | 0.206 | 0.266 | 0.272 | 0.757 | 0.945 | 0.960 | 0.173 |
| Boosted Prob. | 0.370 | 0.361 | 0.701 | 0.993 | 0.999 | 0.029 | 0.434 | 0.410 | 0.761 | 0.986 | 0.994 | 0.049 |
| Platt Scaling | 0.079 | 0.232 | 0.688 | 0.990 | 0.997 | 0.046 | 0.177 | 0.252 | 0.757 | 0.988 | 0.991 | 0.039 |
| P(True) | 0.263 | 0.279 | 0.686 | 0.944 | 0.992 | 0.283 | 0.288 | 0.297 | 0.753 | 0.960 | 0.986 | 0.259 |
| Attention Score | 0.031 | 0.230 | 0.564 | 0.989 | 0.998 | 0.035 | 0.040 | 0.249 | 0.535 | 0.987 | 0.994 | 0.027 |
| Instruction Model | Reasoning Model | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | ECE | Brier | AUROC | P-RB | A-STB | A-SST | ECE | Brier | AUROC | P-RB | A-STB | A-SST |
| Seq. Likelihood | 0.407 | 0.334 | 0.832 | 0.950 | 0.988 | 0.206 | 0.301 | 0.209 | 0.877 | 0.955 | 0.971 | 0.157 |
| Boosted Prob. | 0.584 | 0.531 | 0.857 | 0.982 | 0.996 | 0.068 | 0.537 | 0.427 | 0.882 | 0.968 | 0.980 | 0.091 |
| Platt Scaling | 0.264 | 0.257 | 0.832 | 0.990 | 0.998 | 0.041 | 0.328 | 0.236 | 0.877 | 0.991 | 0.994 | 0.033 |
| P(True) | 0.307 | 0.307 | 0.797 | 0.921 | 0.990 | 0.257 | 0.421 | 0.354 | 0.795 | 0.953 | 0.980 | 0.229 |
| Attention Score | 0.053 | 0.205 | 0.612 | 0.986 | 0.997 | 0.049 | 0.043 | 0.174 | 0.623 | 0.989 | 0.992 | 0.033 |
| Concept | Prompt Name | Prompt |
|---|---|---|
| Answer | A1 | Answer the question, give ONLY the answer, no other words or explanation: |
| A2 | Answer the question, give ONLY the answer without explanation: | |
| A3 | Answer the question, give ONLY the answer: | |
| A4 | Answer the question as short as possible: | |
| A5 | Answer the question with minimal words: | |
| Provide | P1 | Provide an answer for the question, give ONLY the answer, no other words or explanation: |