Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated output. However, this formulation is misaligned with many practical settings where the central question is how likely a model is to solve a query overall. We show that this mismatch results from the stochastic nature of modern LLM decoding, under which single-response correctness fails to reflect underlying model capability. To address this issue, we introduce capability calibration, a new evaluation framework for measuring how well query-level confidence aligns with a model's expected accuracy on individual queries. We formally distinguish capability calibration (CC) from response calibration (RC) and show that the two differ both theoretically and empirically. We further show that CC is better suited than RC to applications like pass@k prediction and inference budget allocation. Finally, we evaluate common confidence estimation methods to understand the practical feasibility of CC.
Figures & tables
Figure 1: (a) Response calibration. Given an input x , a model fθ , and its single sampled output y^ , response calibration calibrates confidence s(x,y^) against correctness C of y^ . (b) Capability calibration. We propose capability calibration, which calibrates confidence s(x,fθ) against expected accuracy μ of fθ ’s output distribution. While response calibration is limited by LLM stochastic decoding, capability calibration directly targets expected accuracy, providing a more reliable and efficient metric for tasks like forecasting test-time scaling performance and resource allocation.
Definition
Calibration Target
Interpretation
Dependence on LM fθ
Response Calibration
Accuracy of y^ given x
How likely is y^ correct?
No. Since y^ is already decoded, the estimation of s(x,y^) is decoupled from the generating model fθ .
Capability Calibration
Expected accuracy of fθ given x
How likely is fθ to answer x correctly?
Yes. The expected accuracy for x depends on fθ ’s capability.
Table 1: Comparison of calibration definitions. Unlike existing response calibration that assesses whether the confidence estimate s aligns with the correctness of one decoded answer y^ , our proposed capability calibration evaluates whether s aligns with the model fθ ’s capability to answer query x .
Figure 2: Divergence of calibration targets. We estimate the Response Calibration (RC) targets C(x,y^) and Capability Calibration (CC) targets Ey^∼fθ(⋅∣x)[C(x,y^)] . Since response stochasticity does not guarantee correctness stochasticity, the experiment result reveals divergence between the two targets: queries where the RC label is 0 exhibit CC values spanning the full [0, 1] range. The same observation for queries with an RC label of 1. This confirms that response-level outcomes do not reflect the model’s true ability to answer a query.
Figure 4
Brier score (↓)
Domain
Factual knowledge
Mathematical reasoning
General exams
Method
Cost
TriviaQA
SimpleQA
GSM8K
MATH
AIME
MMLU
GPQA
Olmo-3-7B-Instruct
Uniform random baseline
N/A
0.2745
0.3133
0.3119
0.2940
0.2462
0.2565
0.2125
Verbalized confidence
L
0.2717
0.2741
0.0444
0.0632
0.2126
0.1556
0.2672
+ isotonic regression
L
0.1840
0.0138
0.0397
0.0501
0.1544
0.1200
0.1205
P(True)
1
0.1698
0.0416
0.1266
0.1392
0.1791
0.2446
0.1520
Table 3: Capability calibration performance of existing confidence estimators with three LLMs on seven datasets. We use bold to denote the best calibrated method, and underline to denote the second best . While the calibration of verbalized confidence and P(True) varies across models, applying isotonic regression yields improvements. Verbalized confidence achieves the best performance for gpt-oss-20b, while Probe is generally the most effective for the other two models.
Figure 4: Cost-performance tradeoff of different methods. We compare inference cost (x-axis, log-scale) against calibration performance (y-axis, 1 - Brier score), where the upper-left corner is the ideal region. See Figure 5 for full results.
Figure 5: Cost-performance tradeoff of different confidence estimation methods with three LLMs on seven datasets. Following Figure 4 , we compare inference cost (x-axis, average response tokens) against calibration performance (y-axis). Probes consistently outperform the random baseline with the lowest cost, while response consistency incurs a cost higher than decoding responses.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Representative empirical results showing the standard error of expected accuracies obtained by different keval .
Figure 7: Brier score of confidence estimators under different keval . Probes are trained by training dataset of TriviaQA ( Joshi et al., 2017 ) , GSM8K ( Cobbe et al., 2021 ) , and MATH ( Hendrycks et al., 2021b ) . Experimental results show that the Brier score is stable at keval=100 for all scenarios.
Model
Temperature
TriviaQA
GSM8K
MMLU
Olmo-3-7B-Instruct
0.6 / 0.3 / 0.15
0.4884 / 0.4933 / 0.4957
0.9331 / 0.9340 / 0.9348
0.7006 / 0.6989 / 0.6989
Qwen3-8B-non-thinking
0.7 / 0.35 / 0.175
0.6040 / 0.6122 / 0.6140
0.9356 / 0.9355 / 0.9361
0.7940 / 0.7979 / 0.7976
gpt-oss-20b
1.0 / 0.5 / 0.25
0.6507 / 0.6625 / 0.6565
0.9576 / 0.9585 / 0.9566
0.8444 / 0.8429 / 0.8384
Appendix
Table 4: Average expected accuracy of three datasets.
TriviaQA
GSM8K
MMLU
Model
Temp. pair
Spearman’s ρ
MSE
Spearman’s ρ
MSE
Spearman’s ρ
MSE
Olmo-3-7B-Instruct
0.6↔0.3
0.9721
0.0048
0.8654
0.0012
0.9763
0.0034
0.6↔0.15
0.9522
0.0106
0.8443
0.0026
0.9654
0.0063
Qwen3-8B-non-thinking
0.7↔0.35
0.9642
0.0052
0.8680
0.0016
0.9596
0.0026
0.7↔0.175
0.9348
0.0106
0.8129
0.0033
0.9404
0.0059
gpt-oss-20b
1.0↔0.5
0.9449
0.0120
0.6734
0.0011
0.8784
0.0035
Appendix
Table 5: Instance-level expected accuracy compared with the recommended temperature (Temp.).
Olmo-3-7B-Instruct
Temp. 0.3
Temp. 0.15
Method
TriviaQA
GSM8K
MMLU
TriviaQA
GSM8K
MMLU
Uniform random baseline
0.2748
0.3143
0.2594
0.2743
0.3205
0.2594
Verbalized confidence
0.2878
0.0483
0.1489
0.2950
0.0480
0.1576
Verbalized confidence+Iso
0.1684
0.0382
0.1276
0.1764
0.0389
0.1346
P(True)
0.1794
0.1306
0.2272
0.1858
0.1317
0.2306
P(True)+Iso
0.1663
0.0410
0.1283
0.1730
0.0416
0.1336
Appendix
Table 6: Brier score of the confidence estimator at different temperatures. Iso denotes isotonic regression.
Model / Dataset
TriviaQA
SimpleQA
GSM8K
MATH-500
AIME
MMLU
GPQA
Olmo-3-7B-Instruct
48.84%
1.92%
93.31%
90.06%
47.29%
70.06%
43.64%
Qwen3-8B
60.40%
3.10%
93.56%
82.96%
20.97%
79.40%
49.65%
gpt-oss-20b
65.07%
3.90%
95.76%
95.66%
69.10%
84.44%
66.97%
Appendix
Table 7: Three LLMs’ mean expected accuracies in seven datasets.
Figure 8: The prompt for verbalized confidence.
Figure 9: The prompt for P(True).
Hyperparameter or Model Info.
Olmo-3-7B-Ins.
Qwen3-8B
gpt-oss-20b
Number of layer activations used
33
37
25
Hidden dimension
4096
4096
2880
Epochs
100
100
100
Batch size
32
32
32
Weight decay
0.01
0.01
0.01
Pooling method
Mean pooling
Mean pooling
Mean pooling
Appendix
Table 8: Hyperparameters or model information for linear probes trained on three LLMs.
Domain
Factual knowledge
Mathematical reasoning
General exams
Method
Cost
TriviaQA
SimpleQA
GSM8K
MATH
AIME
MMLU
GPQA
Olmo-3-7B-Instruct
Consistency ( k=5 )
5 L
0.1228
0.1605
0.0290
0.0373
0.1220
0.1220
0.2462
Consistency ( k=10 )
10 L
0.1121
0.1297
0.0272
0.0383
0.1097
0.1103
0.2246
Consistency ( k=20 )
20 L
0.1046
0.1141
0.0264
0.0350
0.1048
0.1030
0.2079
Qwen3-8B
Appendix
Table 9: Capability calibration Brier scores of Response Consistency across different numbers of samples kc . We use bold to denote the best calibrated method, and underline to denote the second best . Larger kc generally improves performance at the cost of higher estimation overhead.
Brier score (↓)
Domain
Factual knowledge
Mathematical reasoning
General exams
Method
Cost
TriviaQA
SimpleQA
GSM8K
MATH
AIME
MMLU
GPQA
Olmo-3-7B-Instruct
Probe (train on TriviaQA)
<1
0.1101
0.0397
0.1377
0.1547
0.1315
0.1209
0.1230
Probe (train on SimpleQA)
<1
0.3337
0.0130
0.5342
0.4510
0.1894
0.3187
0.1557
Probe (train on GSM8K)
<1
0.2924
0.6038
0.0364
0.0496
0.1622
0.1142
0.1635
Probe (train on MATH)
<1
0.2628
0.4810
0.0382
0.0373
0.1195
0.1271
0.1255
Appendix
Table 10: Capability calibration performance (Brier score) of linear probes trained with different training sets. For probes, we use different colors to indicate in-domain in-distribution , in-domain out-of-distribution , and out-domain performance. We use bold to denote the best calibrated method, and underline to denote the second best . Probe performs best under in-domain in-distribution settings and generalizes reasonably well under in-domain out-distribution settings.
Domain
Factual knowledge
Mathematical reasoning
General exams
Method
Cost
TriviaQA
SimpleQA
GSM8K
MATH
AIME
MMLU
GPQA
Olmo-3-7B-Instruct
Probe (train on TriviaQA)
<1
0.1101
0.0397
0.1377
0.1547
0.1315
0.1209
0.1230
Probe (train on GSM8K)
<1
0.2924
0.6038
0.0364
0.0496
0.1622
0.1142
0.1635
Probe (TriviaQA + GSM8K)
<1
0.1147
0.0396
0.0371
0.0545
0.1127
0.1298
0.1141
Qwen3-8B
Appendix
Table 11: Capability calibration Brier scores of linear probes trained on single and mixed datasets. Results reported by Brier scores (↓). For linear probes trained with different datasets, we use different colors to indicate in-domain in-distribution , in-domain out-of-distribution , and out-domain performance. We use bold to denote the best calibrated method, and underline to denote the second best . Results show that training probes on a mixture of datasets (TriviaQA + GSM8K) generally yields the most robust calibration across both in-distribution and in-domain OOD tasks (e.g., MATH, SimpleQA). However, this benefit is less consistent for out-domain datasets (MMLU, GPQA), where specialized single-dataset probes occasionally maintain an edge.
Spearman’s ρ (↑)
Domain
Factual knowledge
Mathematical reasoning
General exams
Method
Cost
TriviaQA
SimpleQA
GSM8K
MATH
AIME
MMLU
GPQA
Olmo-3-7B-Instruct
Uniform random baseline
N/A
0.0000
0.0000
0.0000
0.0000
0.0000
0.0000
0.0000
Verbalized confidence
L
0.3424
0.0982
0.1293
0.1476
0.4013
0.1709
0.1552
+ isotonic regression
L
0.3443
0.0928
0.1325
0.1576
0.3593
0.1696
0.1694
P(True)
1
0.3944
0.1336
0.2474
0.1862
0.1123
0.3163
0.0211
Appendix
Table 12: Capability calibration performance (Spearman’s ρ ) of different methods with three LLMs on seven datasets. We use bold to denote the best calibrated method, and underline to denote the second best . The findings are consistent with the Brier score result in Table 3 .
ECE ( ↓ )
Domain
Factual knowledge
Mathematical reasoning
General exams
Method
Cost
TriviaQA
SimpleQA
GSM8K
MATH
AIME
MMLU
GPQA
Olmo-3-7B-Instruct
Uniform random baseline
N/A
0.2511
0.4667
0.4365
0.4110
0.2644
0.2905
0.2601
Verbalized confidence
L
0.3222
0.3803
0.0143
0.0309
0.2913
0.1877
0.3714
+ isotonic regression
L
0.0382
0.0141
0.0192
0.0066
0.1597
0.0153
0.0446
P(True)
1
0.0985
0.0847
0.2483
0.2455
0.2062
0.3126
0.1540
Appendix
Table 13: Capability calibration performance (ECE) of different methods with three LLMs on seven datasets. We use bold to denote the best calibrated method, and underline to denote the second best . The findings are consistent with the Brier score results in Table 3 .
Figure 10: Reliability diagrams for the three models.
Method / Model
Olmo-3-7B-Instruct
Qwen3-8B
gpt-oss-20b
Uniform random baseline
0.2745
0.2865
0.2639
Verbalized confidence
0.2717
0.2459
0.1204
Verbalized confidence (few-shot)
0.2802
0.2724
0.1217
P(True)
0.1698
0.3253
0.2435
Probe
0.1101
0.1163
0.1027
Appendix
Table 14: Performance of few-shot verbalized confidence on TriviaQA (left) and GSM8K (right) with reduced offline training costs . Few-shot examples are from the training set of TriviaQA and GSM8K, respectively. Three few-shot examples are provided. Other confidence estimators are listed for comparison. Results are reported by Brier scores ( ↓ ). While few-shot verbalized confidence consistently improves upon the zero-shot baseline on GSM8K, this advantage is not observed consistently on TriviaQA.
Method / Model
Olmo-3-7B-Instruct
Qwen3-8B
gpt-oss-20b
Uniform random baseline
0.2745
0.2865
0.2639
Verbalized confidence
0.2717
0.2459
0.1204
P(True)
0.1698
0.3253
0.2435
Probe ( ktrain=100 )
0.1101
0.1163
0.1027
Probe ( ktrain=10 )
0.1117
0.1083
0.0859
Appendix
Table 15: Performance of linear probes on TriviaQA (left) and GSM8K (right) with reduced offline training costs . Results are reported by Brier scores ( ↓ ). Other confidence estimators are listed for comparison. Results show that linear probes trained with ktrain=10 achieve nearly identical performance to the full setup ( ktrain=100 ) across all evaluated models.
Method
pass@1
pass@4
pass@16
pass@64
Olmo-3-7B-Instruct
Oracle-CC
0.0000
0.0000
0.0001
0.0014
Oracle-RC
0.0967
0.1523
0.2541
0.3395
Probe
0.1092
0.2002
0.1928
0.1577
Qwen3-8B
Oracle-CC
0.0000
0.0000
0.0002
0.0036
Appendix
Table 16: Pass@ k simulation error (MSE) on the AIME dataset. The experiment findings are consistent with the MATH-500 simulation results at Table 2 .
Figure 11: Pass@ k simulation results across 3 models on MATH-500 and AIME datasets. While Oracle-CC can simulate the actual pass@ k near-perfectly, Oracle-RC performs poorly as it is not measuring the models’ capabilities. However, Probe does not fit well in many scenarios.
Figure 12: Budget allocation results across 3 models on MATH-500 and AIME datasets. Oracle confidence consistency reaches the best performance in all scenarios. Verbalized confidence and Probe often outperform the uniform allocation. If the Brier score is good (see Table 3 for the Brier score), the estimated confidence will have better performance in inference budget allocation.
CC choice; RC alt. (%)
TriviaQA
SimpleQA
GSM8K
MATH
AIME
MMLU
GPQA
gpt-oss-20b
P; P
P; P
V; V
V; P (12.3%)
V; P (6.0%)
V; V
V; V
Appendix
Table 17: CC selections and alternative single-response RC selections for gpt-oss-20b. We report an alternative RC selection ( RC alt. ) only when it is selected in at least 5% of trials. P and V denote Probe and Verbalized confidence, respectively.
Data
Selection
Estimator
Pass@1
Pass@4
Pass@16
Pass@64
MATH
CC choice
Verbalized
0.0186
0.0102
0.0067
0.0028
RC alternative
Probe (12.3%)
0.0216
0.0102
0.0068
0.0028
AIME
CC choice
Verbalized
0.0581
0.0717
0.0701
0.0584
RC alternative
Probe (6.0%)
0.0966
0.0968
0.0744
0.0584
Appendix
Table 18: Pass@ k prediction MSE for the CC-selected estimator and the alternative RC selection. Lower is better. RC selection probabilities are shown in parentheses.
Figure 13: Divergence of calibration targets across three models and seven datasets. The data show that the Response Calibration (RC) target C(x,y^) and Capability Calibration (CC) targets Ey^∼fθ(⋅∣x)[C(x,y^)] differs at each model-dataset pair.
Prior work has shown that instruction-tuned large language models (LLMs) are less well calibrated than their base pre-trained counterparts. However, little is known about the frequently used chat template's effect on the calibration of conversational LLMs. In this work, we investigate the mechanisms driving this miscalibration by decoupling the effects of the post-training algorithm and the chat format. We find that, while instruction tuning fundamentally harms calibration, the chat template aggravates the issue through an "ownership bias" -- models are significantly more confident in their own answers than in identical answers provided by a user. Extensive experiments across six recent open-weight LLMs, three benchmarks, and three confidence elicitation methods show that models assign up to 26% higher confidence to their own responses. Leveraging this insight, we propose a simple inference-time strategy: framing the model's answer as user input during confidence elicitation. This approach significantly reduces overconfidence and improves calibration by up to 26% without the need for retraining, narrowing the gap between base and instruction-tuned models.
Mario Sanz-Guerrero, Manuel Mager, Katharina von der Wense
∇Johannes Gutenberg University Mainz, Germany · ♠University of Colorado Boulder, USA
Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful. We argue that this adoption gap is a major obstacle to trustworthy LLM evaluation. Miscalibration causes problems in two distinct areas: at deployment, where overconfident mistakes cause real harm, and inside the research pipeline, where methods like LLM-as-a-judge, synthetic data generation, and active learning rely on calibrated confidence without verifying it. Standard calibration metrics only require two inputs per example: a confidence score and a correctness judgment. Most benchmarks in use today already provide both, meaning calibration can be reported immediately. For open-ended generation, however, defining these two inputs is still an open challenge. We argue that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.
Mario Sanz-Guerrero, Katharina von der Wense
Johannes Gutenberg University Mainz, Germany · University of Colorado Boulder, USA
Existing calibration methods for Large Language Models (LLMs) often overlook a critical dimension of trustworthiness: a model's {\em behavioral robustness} to irrelevant or misleading information. In this paper, we argue that a model's true confidence should reflect its stability under cognitive pressure. We introduce \textsc{CaliDist}, a novel post-hoc calibration approach that directly measures and penalizes a model's susceptibility to distraction. \textsc{CaliDist} quantifies how an LLM's predictions and uncertainty change when its input prompt is perturbed with semantic \textit{distractors}. This stability (or lack thereof) signal is then used to adaptively scale the model's initial confidence score. Our extensive experiments on seven Natural Language Understanding classification benchmarks using six distinct LLMs show that \textsc{CaliDist} consistently achieves lower Expected Calibration Error (ECE) and Brier Score compared with strong baselines. Remarkably, our method reduces the ECE from 23% to 7% on average--a relative improvement of 70%--demonstrating that behavioral stability is a powerful signal for calibration. We make our code and datasets available at github.com/m-anas-j/CaliDist.
Mohammad Anas Jawad, Cornelia Caragea
Computer Science, University of Illinois Chicago, USA.