On Calibration of Large Language Models: From Response To Capability
Organizations: Appier AI Research · National Taiwan University
Abstract
Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on response-level confidence, which estimates the correctness of a single generated output. However, this formulation is misaligned with many practical settings where the central question is how likely a model is to solve a query overall. We show that this mismatch results from the stochastic nature of modern LLM decoding, under which single-response correctness fails to reflect underlying model capability. To address this issue, we introduce capability calibration, a new evaluation framework for measuring how well query-level confidence aligns with a model's expected accuracy on individual queries. We formally distinguish capability calibration (CC) from response calibration (RC) and show that the two differ both theoretically and empirically. We further show that CC is better suited than RC to applications like pass@k prediction and inference budget allocation. Finally, we evaluate common confidence estimation methods to understand the practical feasibility of CC.
Figures & tables
| Definition | Calibration Target | Interpretation | Dependence on LM |
| Response Calibration | Accuracy of given | How likely is correct? | No. Since is already decoded, the estimation of is decoupled from the generating model . |
| Capability Calibration | Expected accuracy of given | How likely is to answer correctly? | Yes. The expected accuracy for depends on ’s capability. |
| Brier score (↓) | Domain | Factual knowledge | Mathematical reasoning | General exams | ||||
| Method | Cost | TriviaQA | SimpleQA | GSM8K | MATH | AIME | MMLU | GPQA |
| Olmo-3-7B-Instruct | ||||||||
| Uniform random baseline | N/A | 0.2745 | 0.3133 | 0.3119 | 0.2940 | 0.2462 | 0.2565 | 0.2125 |
| Verbalized confidence | 0.2717 | 0.2741 | 0.0444 | 0.0632 | 0.2126 | 0.1556 | 0.2672 | |
| + isotonic regression | 0.1840 | 0.0138 | 0.0397 | 0.0501 | 0.1544 | 0.1200 | 0.1205 | |
| P(True) | 1 | 0.1698 | 0.0416 | 0.1266 | 0.1392 | 0.1791 | 0.2446 | 0.1520 |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Temperature | TriviaQA | GSM8K | MMLU |
| Olmo-3-7B-Instruct | 0.6 / 0.3 / 0.15 | 0.4884 / 0.4933 / 0.4957 | 0.9331 / 0.9340 / 0.9348 | 0.7006 / 0.6989 / 0.6989 |
| Qwen3-8B-non-thinking | 0.7 / 0.35 / 0.175 | 0.6040 / 0.6122 / 0.6140 | 0.9356 / 0.9355 / 0.9361 | 0.7940 / 0.7979 / 0.7976 |
| gpt-oss-20b | 1.0 / 0.5 / 0.25 | 0.6507 / 0.6625 / 0.6565 | 0.9576 / 0.9585 / 0.9566 | 0.8444 / 0.8429 / 0.8384 |
| TriviaQA | GSM8K | MMLU | |||||
| Model | Temp. pair | Spearman’s | MSE | Spearman’s | MSE | Spearman’s | MSE |
| Olmo-3-7B-Instruct | 0.9721 | 0.0048 | 0.8654 | 0.0012 | 0.9763 | 0.0034 | |
| 0.9522 | 0.0106 | 0.8443 | 0.0026 | 0.9654 | 0.0063 | ||
| Qwen3-8B-non-thinking | 0.9642 | 0.0052 | 0.8680 | 0.0016 | 0.9596 | 0.0026 | |
| 0.9348 | 0.0106 | 0.8129 | 0.0033 | 0.9404 | 0.0059 | ||
| gpt-oss-20b | 0.9449 | 0.0120 | 0.6734 | 0.0011 | 0.8784 | 0.0035 | |
| Olmo-3-7B-Instruct | Temp. 0.3 | Temp. 0.15 | ||||
| Method | TriviaQA | GSM8K | MMLU | TriviaQA | GSM8K | MMLU |
| Uniform random baseline | 0.2748 | 0.3143 | 0.2594 | 0.2743 | 0.3205 | 0.2594 |
| Verbalized confidence | 0.2878 | 0.0483 | 0.1489 | 0.2950 | 0.0480 | 0.1576 |
| Verbalized confidence+Iso | 0.1684 | 0.0382 | 0.1276 | 0.1764 | 0.0389 | 0.1346 |
| P(True) | 0.1794 | 0.1306 | 0.2272 | 0.1858 | 0.1317 | 0.2306 |
| P(True)+Iso | 0.1663 | 0.0410 | 0.1283 | 0.1730 | 0.0416 | 0.1336 |
| Model / Dataset | TriviaQA | SimpleQA | GSM8K | MATH-500 | AIME | MMLU | GPQA |
| Olmo-3-7B-Instruct | 48.84% | 1.92% | 93.31% | 90.06% | 47.29% | 70.06% | 43.64% |
| Qwen3-8B | 60.40% | 3.10% | 93.56% | 82.96% | 20.97% | 79.40% | 49.65% |
| gpt-oss-20b | 65.07% | 3.90% | 95.76% | 95.66% | 69.10% | 84.44% | 66.97% |
| Hyperparameter or Model Info. | Olmo-3-7B-Ins. | Qwen3-8B | gpt-oss-20b |
| Number of layer activations used | 33 | 37 | 25 |
| Hidden dimension | 4096 | 4096 | 2880 |
| Epochs | 100 | 100 | 100 |
| Batch size | 32 | 32 | 32 |
| Weight decay | 0.01 | 0.01 | 0.01 |
| Pooling method | Mean pooling | Mean pooling | Mean pooling |
| Domain | Factual knowledge | Mathematical reasoning | General exams | |||||
| Method | Cost | TriviaQA | SimpleQA | GSM8K | MATH | AIME | MMLU | GPQA |
| Olmo-3-7B-Instruct | ||||||||
| Consistency ( ) | 5 | 0.1228 | 0.1605 | 0.0290 | 0.0373 | 0.1220 | 0.1220 | 0.2462 |
| Consistency ( ) | 10 | 0.1121 | 0.1297 | 0.0272 | 0.0383 | 0.1097 | 0.1103 | 0.2246 |
| Consistency ( ) | 20 | 0.1046 | 0.1141 | 0.0264 | 0.0350 | 0.1048 | 0.1030 | 0.2079 |
| Qwen3-8B | ||||||||
| Brier score (↓) | Domain | Factual knowledge | Mathematical reasoning | General exams | ||||
| Method | Cost | TriviaQA | SimpleQA | GSM8K | MATH | AIME | MMLU | GPQA |
| Olmo-3-7B-Instruct | ||||||||
| Probe (train on TriviaQA) | 0.1101 | 0.0397 | 0.1377 | 0.1547 | 0.1315 | 0.1209 | 0.1230 | |
| Probe (train on SimpleQA) | 0.3337 | 0.0130 | 0.5342 | 0.4510 | 0.1894 | 0.3187 | 0.1557 | |
| Probe (train on GSM8K) | 0.2924 | 0.6038 | 0.0364 | 0.0496 | 0.1622 | 0.1142 | 0.1635 | |
| Probe (train on MATH) | 0.2628 | 0.4810 | 0.0382 | 0.0373 | 0.1195 | 0.1271 | 0.1255 | |
| Domain | Factual knowledge | Mathematical reasoning | General exams | |||||
| Method | Cost | TriviaQA | SimpleQA | GSM8K | MATH | AIME | MMLU | GPQA |
| Olmo-3-7B-Instruct | ||||||||
| Probe (train on TriviaQA) | 0.1101 | 0.0397 | 0.1377 | 0.1547 | 0.1315 | 0.1209 | 0.1230 | |
| Probe (train on GSM8K) | 0.2924 | 0.6038 | 0.0364 | 0.0496 | 0.1622 | 0.1142 | 0.1635 | |
| Probe (TriviaQA + GSM8K) | 0.1147 | 0.0396 | 0.0371 | 0.0545 | 0.1127 | 0.1298 | 0.1141 | |
| Qwen3-8B | ||||||||
| Spearman’s (↑) | Domain | Factual knowledge | Mathematical reasoning | General exams | ||||
| Method | Cost | TriviaQA | SimpleQA | GSM8K | MATH | AIME | MMLU | GPQA |
| Olmo-3-7B-Instruct | ||||||||
| Uniform random baseline | N/A | 0.0000 | 0.0000 | 0.0000 | 0.0000 | 0.0000 | 0.0000 | 0.0000 |
| Verbalized confidence | 0.3424 | 0.0982 | 0.1293 | 0.1476 | 0.4013 | 0.1709 | 0.1552 | |
| + isotonic regression | 0.3443 | 0.0928 | 0.1325 | 0.1576 | 0.3593 | 0.1696 | 0.1694 | |
| P(True) | 1 | 0.3944 | 0.1336 | 0.2474 | 0.1862 | 0.1123 | 0.3163 | 0.0211 |
| ECE ( ) | Domain | Factual knowledge | Mathematical reasoning | General exams | ||||
| Method | Cost | TriviaQA | SimpleQA | GSM8K | MATH | AIME | MMLU | GPQA |
| Olmo-3-7B-Instruct | ||||||||
| Uniform random baseline | N/A | 0.2511 | 0.4667 | 0.4365 | 0.4110 | 0.2644 | 0.2905 | 0.2601 |
| Verbalized confidence | 0.3222 | 0.3803 | 0.0143 | 0.0309 | 0.2913 | 0.1877 | 0.3714 | |
| + isotonic regression | 0.0382 | 0.0141 | 0.0192 | 0.0066 | 0.1597 | 0.0153 | 0.0446 | |
| P(True) | 1 | 0.0985 | 0.0847 | 0.2483 | 0.2455 | 0.2062 | 0.3126 | 0.1540 |
| Method / Model | Olmo-3-7B-Instruct | Qwen3-8B | gpt-oss-20b |
| Uniform random baseline | 0.2745 | 0.2865 | 0.2639 |
| Verbalized confidence | 0.2717 | 0.2459 | 0.1204 |
| Verbalized confidence (few-shot) | 0.2802 | 0.2724 | 0.1217 |
| P(True) | 0.1698 | 0.3253 | 0.2435 |
| Probe | 0.1101 | 0.1163 | 0.1027 |
| Method / Model | Olmo-3-7B-Instruct | Qwen3-8B | gpt-oss-20b |
| Uniform random baseline | 0.2745 | 0.2865 | 0.2639 |
| Verbalized confidence | 0.2717 | 0.2459 | 0.1204 |
| P(True) | 0.1698 | 0.3253 | 0.2435 |
| Probe ( ) | 0.1101 | 0.1163 | 0.1027 |
| Probe ( ) | 0.1117 | 0.1083 | 0.0859 |
| Method | pass@1 | pass@4 | pass@16 | pass@64 |
| Olmo-3-7B-Instruct | ||||
| Oracle-CC | 0.0000 | 0.0000 | 0.0001 | 0.0014 |
| Oracle-RC | 0.0967 | 0.1523 | 0.2541 | 0.3395 |
| Probe | 0.1092 | 0.2002 | 0.1928 | 0.1577 |
| Qwen3-8B | ||||
| Oracle-CC | 0.0000 | 0.0000 | 0.0002 | 0.0036 |
| CC choice; RC alt. (%) | TriviaQA | SimpleQA | GSM8K | MATH | AIME | MMLU | GPQA |
| gpt-oss-20b | P; P | P; P | V; V | V; P (12.3%) | V; P (6.0%) | V; V | V; V |
| Data | Selection | Estimator | Pass@1 | Pass@4 | Pass@16 | Pass@64 |
| MATH | CC choice | Verbalized | 0.0186 | 0.0102 | 0.0067 | 0.0028 |
| RC alternative | Probe (12.3%) | 0.0216 | 0.0102 | 0.0068 | 0.0028 | |
| AIME | CC choice | Verbalized | 0.0581 | 0.0717 | 0.0701 | 0.0584 |
| RC alternative | Probe (6.0%) | 0.0966 | 0.0968 | 0.0744 | 0.0584 |