As AI systems move from static repositories to agents that are capable of continual adaptation and learning, maintaining their trustworthiness means equipping the models backing them with the ability to produce confidence estimates that dynamically reflect their changing skills and knowledge. We introduce the problem of persistent calibration, which requires a confidence estimator to faithfully reflect the knowledge contained in a model as that knowledge changes, without recurring supervision. We operationalize this by examining persistent calibration across checkpoints of open models, asking whether confidence estimators trained on earlier checkpoints can generalize to later ones. Specifically, we aim to shed light on whether confidence is dependent on knowledge, a question with implications for the reliability of confidence estimates. To measure this relationship, we define and evaluate calibration on knowledge contrast sets: subsets containing questions that one checkpoint answers correctly and another checkpoint answers incorrectly, reflecting a change in knowledge. We show that both inference-time and fine-tuning methods fall short on contrast-set calibration compared to oracle methods trained on future checkpoints, even for methods that are well-calibrated on the full dataset. We provide evidence for the hypothesis that persistent calibration is challenging because there is a vast space of possible confidence functions that are well-calibrated on a given checkpoint, out of which only some rely on meta-knowledge features that would generalize to other checkpoints. Towards improving contrast-set calibration, we show that multi-checkpoint training helps, suggesting an avenue for identifying confidence features that remain robust across changing knowledge.
Figures & tables
Figure 1: (Left) Our setup for evaluating persistent calibration. During training, the model encounters the training data instance “Paris is the capital of France” , causing a change in knowledge between checkpoints {\color[rgb]{0.8477,0.1055,0.375}\text{C}_{\text{2}}} and {\color[rgb]{0.1172,0.5352,0.8984}\text{C}_{\text{3}}} . A confidence estimator that is faithful to the model’s evolving knowledge (Hypothesis 2) should produce a higher confidence on the checkpoint {\color[rgb]{0.1172,0.5352,0.8984}\text{C}_{\text{3}}} with increased knowledge; however, an estimator could instead capture Hypothesis 1, which performs equally well on the training distribution {\color[rgb]{0.8477,0.1055,0.375}\text{C}_{\text{2}}} but fails to generalize to {\color[rgb]{0.1172,0.5352,0.8984}\text{C}_{\text{3}}} . (Right) To measure the dependence of confidence on knowledge, we evaluate calibration on a contrast set consisting of questions that one checkpoint answers correctly and another checkpoint answers incorrectly, testing whether a change in knowledge is accompanied by a change in confidence.
Olmo 3 7B
Marin 8B
Olmo 3 32B
Contrast
Full
Contrast
Full
Contrast
Full
Method
Δ0b↑
AUC ↑
AUC ↑
Δ0b↑
AUC ↑
AUC ↑
Δ0b↑
AUC ↑
AUC ↑
TriviaQA
End-correct baseline
0.000
0.627
0.530
0.000
0.783
0.547
0.000
0.712
0.548
Self-consistency
0.297
0.733
0.884
0.263
0.823
0.897
0.362
0.808
0.903
Non-oracle training
0.333
0.748
0.880
0.227
0.806
0.834
0.285
0.688
0.860
Table 1: SC and non-oracle training consistently fall short on contrast-set calibration compared to oracle training, even when full-set calibration is saturated. Results in Section 3 are averaged over 3 seeds. In each column, we bold the best result and underline results not significantly worse under a paired test ( α=0.05 ; see Section A.8 for tests); tables in Section 3 share this text styling. See Section B.2 for all evaluation metrics and Jeopardy results.
Olmo 3 7B
Marin 8B
Olmo 3 32B
Contrast
Full
Contrast
Full
Contrast
Full
Method
Δ0b↑
AUC ↑
AUC ↑
Δ0b↑
AUC ↑
AUC ↑
Δ0b↑
AUC ↑
AUC ↑
Self-consistency
0.297
0.733
0.884
0.263
0.823
0.897
0.362
0.808
0.903
└ Surrogate
0.211
0.637
0.859
0.277
0.695
0.871
0.288
0.688
0.883
└ Copy
0.000
0.500
0.832
0.000
0.500
0.829
0.000
0.500
0.846
Non-oracle training
0.333
0.748
0.880
0.227
0.806
0.834
0.285
0.688
0.860
Table 2: Comparing SC and non-oracle training with ablations controlling for knowledge already present in an earlier checkpoint. Beating these ablations is necessary for a nontrivial demonstration of persistent calibration. SC generally beats its ablations, while non-oracle training has mixed results. Results on TriviaQA; see Section B.2 for all metrics and Jeopardy results.
Olmo 3 7B
Marin 8B
Olmo 3 32B
Contrast
Full
Contrast
Full
Contrast
Full
Method
Δ0b↑
AUC ↑
AUC ↑
Δ0b↑
AUC ↑
AUC ↑
Δ0b↑
AUC ↑
AUC ↑
Non-oracle
Multi-ckpt
0.333
0.748
0.880
0.227
0.806
0.834
0.285
0.688
0.860
Single-ckpt
0.332
0.750
0.876
0.221
0.642
0.783
0.251
0.673
0.849
Oracle
Table 3: Multi-checkpoint training improves contrast-set and full-set calibration. Results on TriviaQA; see Section B.3 for Jeopardy.
Contrast
Full
Method
Δ0b↑
Δ0↑
Δb↑
Δ↑
AUC ↑
BS ↓
ECE ↓
AUC ↑
BS ↓
ECE ↓
Multi-ckpt
0.379
0.424
0.255
0.285
0.795
0.185
0.061
0.890
0.131
0.032
└ Single-ckpt data
0.373
0.368
0.252
0.258
0.768
0.197
0.096
0.887
0.134
0.044
Resp-ckpt
0.351
0.396
0.241
0.274
0.778
0.193
0.070
0.884
0.133
0.030
└ Pooled data
0.360
0.357
0.248
0.251
0.757
0.203
0.104
0.887
0.132
0.035
Table 4: Ablating factors contributing to the gains of multi-checkpoint training (oracle). Single-checkpointdata uses only one checkpoint’s answers for multi-checkpoint training. Pooleddata trains each checkpoint’s adapter on the same data as (unablated) multi-checkpoint training. Multi-checkpoint training beats both, meaning that both the training data and multiple checkpoint views contribute. Results on TriviaQA, Olmo 3 7B; see Section B.3 for Jeopardy.
Figure 2: GCM’s headroom on contrast-set AUC grows for larger models where the GCM backbone (Qwen3-8B) knows the answer less often. Thus, a static external verifier does not adapt to a model’s capability, motivating meta-knowledge for persistent calibration.
Figure 3: (Left) Transfer between pairs of training and eval checkpoints (Olmo 3 7B, TriviaQA; Jeopardy in Section B.6 ). Forward transfer (top right) is strong, while backward transfer (bottom left) is weak. (Right) Transfer of multi-checkpoint training, where the second checkpoint gets random labels. The second-checkpoint AUC generally plummets, at no penalty to the first checkpoint.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
19th Century America
3-Letter Words
4-Letter Words
American History
American Literature
Americana
Animals
Annual Events
Architecture
Around The World
Art
Art & Artists
Astronomy
Authors
Awards
Ballet
Biology
Bodies Of Water
Books & Authors
Business & Industry
Classical Music
Appendix
Table 5: Jeopardy categories included in our evaluation.
Type
Question
Reference answer(s)
Factoid
What is the main use of ETD fragmentation?
Analysis of intact proteins
Factoid
Which disease can be treated with Delamanid?
tuberculosis
List
List approved radioprotective compounds. Just output one of the answers.
[[‘amifostine’], [‘palifermin’]]
List
List the 5 different human immunoglobulin heavy chains. Just output one of the answers.
Table 6: BioASQ example questions. For list questions, an answer matching any of the reference answers is accepted.
Accuracy (%)
Short name
Checkpoint ID
Training data (trillions of tokens)
TriviaQA
Jeopardy
BioASQ
Olmo 3 7B
10%
stage1-step141000
0.59
51.9
64.8
40.2
20%
stage1-step283000
1.19
55.5
68.5
41.7
30%
stage1-step424000
1.78
57.8
70.7
44.0
40%
stage1-step566000
2.37
59.5
72.1
44.5
Appendix
Table 7: Cumulative training data and performance for the checkpoints used in our experiments.
Olmo 3 7B
40%
50%
90%
100%
40%
–
642 (354)
771 (516)
753 (527)
50%
642 (288)
–
715 (455)
697 (466)
90%
771 (255)
715 (260)
–
538 (289)
100%
753 (226)
697 (231)
538 (249)
–
Appendix
Table 8: TriviaQA contrast-set sizes for each pair of evaluation checkpoints. The full-set size is 4960. In a cell, the first number is the total size of the contrast set, and the second number is the number of questions that the row checkpoint gets incorrect and the column checkpoint gets correct. Contrast-set sizes that are grayed out are too small (fewer than 500 questions) and are not used ( Section A.3 ).
Olmo 3 7B
40%
50%
90%
100%
40%
–
738 (406)
767 (526)
796 (559)
50%
738 (332)
–
695 (453)
720 (484)
90%
767 (241)
695 (242)
–
517 (277)
100%
796 (237)
720 (236)
517 (240)
–
Appendix
Table 9: Jeopardy contrast-set sizes for each pair of evaluation checkpoints. The full-set size is 4998. In a cell, the first number is the total size of the contrast set, and the second number is the number of questions that the row checkpoint gets incorrect and the column checkpoint gets correct. Contrast-set sizes that are grayed out are too small (fewer than 500 questions) and are not used ( Section A.3 ).
Olmo 3 7B
40%
50%
90%
100%
40%
–
457 (273)
490 (312)
475 (328)
50%
457 (184)
–
483 (264)
466 (279)
90%
490 (178)
483 (219)
–
389 (218)
100%
475 (147)
466 (187)
389 (171)
–
Appendix
Table 10: BioASQ contrast-set sizes for each pair of evaluation checkpoints. The full-set size is 2719. In a cell, the first number is the total size of the contrast set, and the second number is the number of questions that the row checkpoint gets incorrect and the column checkpoint gets correct. Contrast-set sizes that are grayed out are too small (fewer than 300 questions) and are not used ( Section A.3 ).
Contrast
Full
Method
Δ0b↑
Δ0↑
Δb↑
Δ↑
AUC ↑
BS ↓
ECE ↓
AUC ↑
BS ↓
ECE ↓
Olmo 3 7B, TriviaQA
End-correct baseline
0.000
0.254
0.000
0.236
0.627
0.355
0.340
0.530
0.370
0.308
Self-consistency
0.297
0.329
0.190
0.212
0.733
0.212
0.084
0.884
0.140
0.069
└ Post-hoc
0.298
0.330
0.210
0.234
0.734
0.215
0.097
0.884
0.133
0.019
Non-oracle training
0.333
0.324
0.236
0.239
0.748
0.206
0.104
0.880
0.136
0.032
Appendix
Table 11: Expanding on Table 1 , we consistently find that SC and non-oracle training fall short on contrast-set calibration compared to oracle training, even when full-set calibration is saturated. Notably, even after applying post-hoc calibration to SC (isotonic regression via the pool adjacent violators algorithm fitted on the train set of the non-oracle checkpoints, applied with linear interpolation), full-set ECE becomes negligible and yet contrast-set AUC does not improve.
Contrast
Full
Method
Δ0b↑
Δ0↑
Δb↑
Δ↑
AUC ↑
BS ↓
ECE ↓
AUC ↑
BS ↓
ECE ↓
Olmo 3 7B, TriviaQA
End-correct baseline
0.000
0.254
0.000
0.080
0.627
0.230
0.000
0.530
0.370
0.308
Self-consistency
0.297
0.329
0.161
0.180
0.739
0.205
0.080
0.884
0.140
0.069
└ Post-hoc
0.298
0.330
0.163
0.181
0.740
0.205
0.080
0.884
0.133
0.019
Non-oracle training
0.333
0.324
0.211
0.219
0.759
0.196
0.092
0.880
0.136
0.032
Appendix
Table 12: Analogous setup to Table 11 , but using a fitted normalization function ( Section B.1 ) instead of Eq. 1 . The two tables share the finding that SC and non-oracle training underperform oracle training on contrast-set calibration, even when full-set calibration is saturated.
Contrast
Full
Method
Δ0b↑
Δ0↑
Δb↑
Δ↑
AUC ↑
BS ↓
ECE ↓
AUC ↑
BS ↓
ECE ↓
Olmo 3 7B, TriviaQA
Self-consistency
0.297
0.329
0.190
0.212
0.733
0.212
0.084
0.884
0.140
0.069
└ Surrogate
0.211
0.191
0.204
0.184
0.637
0.304
0.208
0.859
0.163
0.094
└ Copy
0.000
0.000
0.000
0.000
0.500
0.250
0.127
0.832
0.161
0.033
Non-oracle training
0.333
0.324
0.236
0.239
0.748
0.206
0.104
0.880
0.136
0.032
Appendix
Table 13: Expanded version of Table 2 , for both TriviaQA and Jeopardy. A confidence estimator must outperform its ablations for a nontrivial demonstration of persistent calibration. Based on contrast-set AUC, SC outperforms its ablations in 5/6 cases, while non-oracle training outperforms its ablations in only 3/6 cases.
Contrast
Full
Method
Δ0b↑
Δ0↑
Δb↑
Δ↑
AUC ↑
BS ↓
ECE ↓
AUC ↑
BS ↓
ECE ↓
Non-oracle
Olmo 3 7B
Multi-ckpt, ncand =2
0.333
0.324
0.236
0.239
0.748
0.206
0.104
0.880
0.136
0.032
Multi-ckpt, ncand =1
0.332
0.335
0.223
0.233
0.749
0.204
0.094
0.877
0.137
0.032
Single-ckpt, ncand =2
0.332
0.340
0.216
0.227
0.750
0.203
0.093
0.876
0.140
0.048
Single-ckpt, ncand =1
0.314
0.339
0.208
0.225
0.746
0.206
0.086
0.870
0.146
0.063
Appendix
Table 14: Multi-checkpoint and multi-answer training often yields substantial gains. TriviaQA results; see Table 15 for Jeopardy.
Contrast
Full
Method
Δ0b↑
Δ0↑
Δb↑
Δ↑
AUC ↑
BS ↓
ECE ↓
AUC ↑
BS ↓
ECE ↓
Non-oracle
Olmo 3 7B
Multi-ckpt, ncand =2
0.476
0.476
0.360
0.365
0.821
0.177
0.094
0.871
0.119
0.038
Multi-ckpt, ncand =1
0.438
0.438
0.327
0.328
0.802
0.186
0.103
0.863
0.124
0.042
Single-ckpt, ncand =2
0.447
0.457
0.318
0.327
0.808
0.181
0.094
0.860
0.124
0.039
Single-ckpt, ncand =1
0.418
0.441
0.278
0.294
0.796
0.185
0.090
0.851
0.130
0.050
Appendix
Table 15: Multi-checkpoint and multi-answer training often yields substantial gains. Jeopardy results; see Table 14 for TriviaQA.
Contrast
Full
Method
Δ0b↑
Δ0↑
Δb↑
Δ↑
AUC ↑
BS ↓
ECE ↓
AUC ↑
BS ↓
ECE ↓
Multi-ckpt
0.506
0.536
0.379
0.402
0.855
0.156
0.057
0.888
0.111
0.020
└ Single-ckpt data
0.517
0.517
0.382
0.383
0.841
0.164
0.087
0.885
0.115
0.041
Resp-ckpt
0.471
0.485
0.352
0.364
0.826
0.172
0.079
0.877
0.116
0.028
└ Pooled data
0.506
0.499
0.380
0.376
0.833
0.169
0.094
0.884
0.113
0.027
Appendix
Table 16: Counterpart to Table 4 (TriviaQA) for Jeopardy. Multi-checkpoint training continues to outperform its two ablations, although the single-checkpoint data ablation is able to match the unablated multi-checkpoint training on the class-balanced contrast-set metrics Δ0b,Δb , suggesting that access to multiple checkpoints’ weights accounts for most of the performance gain of multi-checkpoint training.
Contrast
Full
Method
Δ0b↑
Δ0↑
Δb↑
Δ↑
AUC ↑
BS ↓
ECE ↓
AUC ↑
BS ↓
ECE ↓
Olmo 3 7B, TriviaQA
GCM
0.446
0.448
0.331
0.335
0.808
0.182
0.088
0.869
0.191
0.181
└ Post-hoc
0.446
0.448
0.333
0.337
0.812
0.179
0.085
0.869
0.142
0.035
Oracle training
0.379
0.424
0.255
0.285
0.795
0.185
0.061
0.890
0.131
0.032
Olmo 3 7B, Jeopardy
Appendix
Table 17: GCMs underperform oracle training, especially for larger models, highlighting that a static external verifier does not adapt to a model’s capability, motivating meta-knowledge for persistent calibration. This holds even with post-hoc calibration on GCM; see Table 11 for a description.
Contrast
Full
Method
Δ0b↑
Δ0↑
Δb↑
Δ↑
AUC ↑
BS ↓
ECE ↓
AUC ↑
BS ↓
ECE ↓
Olmo 3 7B, TriviaQA → BioASQ
Self-consistency
0.122
0.137
0.075
0.083
0.602
0.260
0.121
0.738
0.216
0.087
Non-oracle training
0.223
0.239
0.149
0.163
0.680
0.234
0.106
0.805
0.189
0.090
Oracle training
0.211
0.256
0.131
0.155
0.691
0.223
0.072
0.805
0.184
0.054
Olmo 3 7B, Jeopardy → BioASQ
Appendix
Table 18: Transfer from TriviaQA or Jeopardy to BioASQ. The advantage of oracle training over other methods does not hold under domain transfer, consistent with prior work on the domain-specificity of correctness signals ( Sky et al., 2024 ) .
Layer
Mean ρ
Median ρ
1
0.920
0.929
2
0.914
0.929
3
0.631
0.750
4
0.507
0.536
5
0.533
0.714
6
0.535
0.750
Appendix
Table 19: Variance analysis ( Section B.7 ). Spearman correlations between checkpoint index and adapter variance for Olmo 3 7B on TriviaQA. For each question and layer, we compute the total variance (i.e. trace of the covariance matrix) of the model’s hidden states across 10 confidence adapter training seeds, and we compute the correlation between this total variance and checkpoint index. We find that the mean/median correlation across questions is positive for all layers, suggesting that later checkpoints tend to have a larger space of viable confidence features that confidence adapters can pick up on.
Figure 4: (Left) Transfer between pairs of training and eval checkpoints (Olmo 3 7B, Jeopardy, see Fig. 3 for TriviaQA). Forward transfer (top right) is strong, while backward transfer (bottom left) is weak. (Right) Transfer of multi-checkpoint training, where the second checkpoint gets random labels. The second-checkpoint AUC generally plummets, at no penalty to the first checkpoint.
When a model knows when it does not know, many possibilities emerge. The first question is how to enable a model to recognize that it does not know. A promising approach is to use confidence, computed from the model's internal signals, to reflect its ignorance. Prior work in specific domains has shown that calibration can provide reliable confidence estimates. In this work, we propose a simple, effective, and universal training-free method that applies to both vision and language models, performing model calibration, cascading, and data cleaning to better exploit a model's ability to recognize when it does not know. We first highlight two key empirical observations: higher confidence corresponds to higher accuracy within a single model, and models calibrated on the validation set remain calibrated on a held-out test set. These findings empirically establish the reliability and comparability of calibrated confidence. Building on this, we introduce two applications: (1) model cascading with calibrated advantage routing and (2) data cleaning based on model ensemble. Using the routing signal derived from the comparability of calibrated confidences, we cascade large and small models to improve efficiency with almost no compromise in accuracy, and we further cascade two models of comparable scale to achieve performance beyond either model alone. Leveraging multiple experts and their calibrated confidences, we design a simple yet effective data-cleaning method that balances precision and detection rate to identify mislabeled samples in ImageNet and Massive Multitask Language Understanding (MMLU) datasets. Our results demonstrate that enabling models to recognize when they do not know is a practical step toward more efficient, reliable, and trustworthy AI.
Chenjie Hao, Weyl Lu, Yuko Ishiwaka +3
University of California, Davis · SoftBank Corp. · Aizip
Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success. Existing methods typically elicit confidence only once: either before thinking or after answering. We argue that confidence in reasoning models is state-dependent: before thinking, confidence should estimate the chance of the model correctly solving the prompt, while after thinking it should predict whether the realized answer is likely to be correct. This distinction determines the appropriate supervision target: prompt-level success should supervise confidence estimates made after seeing the prompt, while individual answer-level correctness should supervise confidence estimates made after answering. We introduce CALIBER (Calibration Before and After Reasoning), which elicits both estimates and supervises each with the target matched to its information state. Under this unified protocol, CALIBER reduces Expected Calibration Error (ECE) by 52.5% over the strongest single-confidence baseline on BigMathDigits for the 7B model, while achieving the best Brier score and AUROC, and remains within 2.1 points of the best accuracy. Further, on a larger 30B model, CALIBER achieves the best ECE on BigMathDigits while remaining competitive in Brier score and AUROC. Out of distribution, it achieves the best ECE and Brier score on GPQA and TriviaQA, and remains competitive on SimpleQA. Ablations further show that this position-target alignment is most beneficial under distribution shift where it consistently reduces calibration error across all out-of-distribution benchmarks.
Continual learning for large language models is typically evaluated through accuracy retention under sequential fine-tuning. We argue that this perspective is incomplete, because uncertainty reliability can degrade earlier and more sharply than top-1 performance. We study this empirically by measuring conformal coverage and calibration error on sequentially fine-tuned models across three model families and eight task sequences drawn primarily from classification and multiple-choice benchmarks. Across the classification-style settings we study, coverage loss exceeds accuracy loss by a factor of roughly 3.4×±0.5× on average across seeds; in the most pronounced case, coverage drops from 0.92 to 0.61, while accuracy remains within three points of baseline. Standard continual-learning methods that preserve accuracy do not automatically preserve coverage, and naive calibration baselines recover only part of the gap. We propose calibration replay, a lightweight post-hoc procedure that maintains a task-specific held-out buffer and refits a task-specific conformal threshold under the current model after each update. It adds no training-time gradient cost, uses less than one percent of the memory of ordinary experience replay, and typically restores coverage to within two points of nominal at buffer size m=200. We accompany the empirical study with a drift decomposition, a finite-sample recovery theorem showing exact conformal validity under exchangeability, and a mixture-validity proposition explaining why pooled thresholds do not suffice. Our guarantees are stated for classification-style tasks with task-specific buffers; extensions to open-ended generation are exploratory.
Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
Department of Computer Science, Iowa State University · Department of Civil, Construction & Environmental Engineering, Iowa State University