Accurately assessing model confidence is essential for deploying large language models (LLMs) in mission-critical factual domains. While retrieval-augmented generation (RAG) is widely adopted to improve grounding, confidence calibration in RAG settings remains poorly understood. We conduct a systematic study across four benchmarks, revealing that LLMs exhibit poor calibration performance especially when noisy contexts are retrieved. Specifically, contradictory or irrelevant evidence tends to exacerbate the model's overconfidence issue. To address this, we propose NOVA Rules (NOise-Aware Verbal Confidence CAlibration Rules) to provide a principled foundation for resolving overconfidence under noise. We further design NOVA, a noise-aware calibration framework that synthesizes supervision from ~2K HotpotQA examples guided by these rules. By performing supervised fine-tuning (SFT) with this data, NOVA equips models with intrinsic noise awareness without relying on stronger teacher models. Empirical results show that NOVA yields substantial gains, improving ECE scores by 10.9% in-domain and 8.0% out-of-domain. By bridging the gap between retrieval noise and verbal calibration, NOVA paves the way for both accurate and epistemically reliable LLMs.
Figures & tables
Figure 1: An illustrative example of model responses before and after NOVA . By explicitly training the model to assess passage and group level utility prior to answering, NOVA enables more reliable confidence expression under noisy retrieval. The performance plots report results on NQ for Llama-3.1-8B-Instruct and DeepSeek-R1-Distill-Llama-8B , where SFT corresponds to the Label-only SFT setting in Table 2 , and illustrate how NOVA promotes more transparent and grounded human–computer interaction in real-world scenarios.
Method
StrategyQA
HotpotQA
NQ
Bamboogle
Average
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Llama-3.1-8B-Instruct
BM25 (CoT)
0.205
0.485
0.496
0.552
0.369
0.688
0.566
0.557
0.409
0.571
Contriever (CoT)
0.167
0.550
0.585
0.476
0.347
0.649
0.592
0.535
0.423
0.552
Qwen2.5-7B-Instruct
BM25 (CoT)
0.190
0.620
0.439
0.683
0.473
0.747
0.650
0.670
0.438
0.680
Table 1: Real-world RAG results across four datasets demonstrate consistently poor calibration. Notably, average ECE exceeds 0.4, highlighting severe misalignment between verbal confidence and empirical correctness.
Figure 2: Calibration performance of Llama-3.1-8B-Instruct and DeepSeek-R1-Distill-Llama-8B on NQ and Bamboogle under controlled noise settings. The plots display ECE, AUROC, and Average Confidence across four retrieval settings: Gold-only , Gold+Irrelevant (Irr), Gold+Relevant (Rel), and Gold+Counterfactual (Cf). Results show that introducing noise, particularly counterfactual passages, substantially degrades calibration performance.
Figure 3: Overview of the NOVA data pipeline, consisting of three stages: RAG Passage Construction, Training Response Generation, and Multi-stage Data Filtering. In Training Response Generation, the model takes a query q and retrieved passages P with k=3 as input (Input: Q+3P), generates passage-level and group-level judgments Jp,Jg (P Type), then predicts the answer a^ (A) and verbal confidence c^ (C). The pipeline ultimately yields 2K high-quality trajectories for fine-tuning.
Method
StrategyQA
HotpotQA
NQ
Bamboogle
Average
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Llama-3.1-8B-Instruct
Vanilla
0.396
0.602
0.460
0.605
0.465
0.577
0.324
0.636
0.411
0.605
CoT
0.354
0.555
0.444
0.645
0.423
0.611
0.288
0.552
0.377
0.591
Noise-aware
0.376
0.615
0.309
0.642
0.351
0.618
0.217
0.793
0.314
0.667
Ensemble
0.370
0.609
0.397
0.650
0.428
0.619
0.214
0.713
0.352
0.648
Table 2: Calibration performance of various models on four datasets. Scores in bold indicate the best performance, while underlined scores denote the second-best. Results show that NOVA substantially improves calibration and consistently outperforms several baselines, without sacrificing accuracy, as evidenced in Appendix B.8 .
NQ
Bamboogle
Average
NQ
Bamboogle
Average
Method
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Method
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Llama-3.1-8B-Instruct
DeepSeek-R1-Distill-Llama-8B
Vanilla
0.371
0.645
0.212
0.633
0.292
0.639
Vanilla
0.376
0.625
0.154
0.671
0.265
0.648
CoT
0.352
0.670
0.199
0.579
0.276
0.625
CoT
0.373
0.621
0.203
0.633
0.288
0.627
Noise-aware
0.289
0.667
0.140
0.806
0.215
0.737
Noise-aware
0.290
0.605
0.153
0.658
0.222
0.632
Ensemble
0.334
0.693
0.173
0.680
0.254
0.687
Ensemble
0.351
0.590
0.143
0.711
0.247
0.651
Table 3: Out-of-Distribution (OOD) results with 5 passage per query on the NQ and Bamboogle datasets, demonstrating that NOVA maintains robust calibration performance and consistently outperforms several strong baselines even when facing varying amounts of retrieved context in unseen scenarios. More generalization test is in Appendix B.1 .
Average
Average
Method
ECE ↓
AUROC ↑
Method
ECE ↓
AUROC ↑
Llama-3.1-8B-Instruct
DeepSeek-R1-Distill-Llama-8B
Vanilla
0.472
0.616
Vanilla
0.429
0.684
CoT
0.423
0.552
CoT
0.454
0.672
Noise-aware
0.318
0.655
Noise-aware
0.409
0.633
Ensemble
0.354
0.620
Ensemble
0.505
0.672
Table 4: Average calibration performance of four models across four datasets with Contriever. Full results are in Appendix B.4 . Bold and underline mark the best and second-best scores within each model.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
# Questions
Confidence Interval
HotpotQA
800
± 0.0347
StrategyQA
800
± 0.0347
NQ
800
± 0.0347
Bamboogle
150
± 0.0800
Appendix
Table 5: Dataset statistics and 95% confidence intervals of ECE and AUROC.
Hyperparameter
BM25
Contriever
Top- K Retrieval
5
5
Reranker
No
No
Model Specifics
Architecture
Sparse (Probabilistic)
Dense (Bi-Encoder)
Embedding Model
N/A
facebook/ contriever
Max Input Length
N/A
256 tokens
Appendix
Table 6: Retrieval Hyperparameters.
Category
Sub-category
Definition
Counterfactual
—
Passages that are semantically relevant to the question but directly contradict the ground truth answer. They provide specific, plausible-sounding information that supports an incorrect alternative answer.
Entity-relevant
Passages that mention the correct entities in the question but only provide partial, tangential, or incomplete factual information, without containing the evidence needed to answer the question.
Relevant Noise
Relation-relevant
Passages that capture the type of relations required by the question but do not involve the queried entities, thereby providing misleading or insufficient evidence.
Theme-relevant
Passages that are topically aligned with the question and provide high-level background or contextual information, but do not contain entity-level or relation-level facts necessary for answering.
Irrelevant Noise
—
Passages that have little to no semantic relation to the question. They are from unrelated topics or domains and provide no useful information for answering.
Appendix
Table 7: Definitions of noise passages for retrieval-augmented question answering. Relevant noise is further categorized into entity-level, relation-level, and theme-level noise to better simulate real-world conditions.
Model
Total
Kept Responses
(1) Format
(2) Passage
(3) Rule
(4) Alignment
(5) Common
(6) Balance
Judgment
Following
IDs
DS-R1-Llama
96000
85723
39008
34403
5211
2801
1945
DS-R1-Qwen
96000
88201
28481
24586
4611
2801
1945
Llama-3.1
96000
78200
35255
28790
4895
2801
1945
Qwen-2.5
96000
94898
31065
26221
3609
2801
1945
Appendix
Table 8: Training data statistics: This table shows the number of training data left after each filtering step. (1) Format: retains only samples from which a valid answer, a confidence score, and intermediate passage judgments can be successfully extracted. (2) Passage judgment: filters out samples containing incorrect assessments of the retrieved passages. (3) Rule following: filters for samples that have a explicit reasoning process for rule following. (4) Alignment: for each query, selects the final response that minimizes the instance-level Brier Score. (5) Common IDs: retains only samples with question IDs common across all models. (6) Balance: balances the 3 groups (counterfactual, consistent, irrelevant) by downsampling consistent to match irrelevant. Model name abbreviations: DS-R1-Llama : DeepSeek-R1-Distill-Llama-8B; DS-R1-Qwen : DeepSeek-R1-Distill-Qwen-7B; Llama-3.1 : Llama-3.1-8B-Instruct; Qwen-2.5 : Qwen2.5-7B-Instruct.
Figure 4: Reliability Diagram for HotpotQA: comparison of CoT prompt with base model (upper row) and SFT models (lower row). Each subplot displays accuracy v.s. confidence, with the diagonal dashed line representing perfect calibration.
Retriever
Prompt Type
StrategyQA
HotpotQA
NQ
Bamboogle
Average
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Llama-3.1-8B-Instruct
BM25
Vanilla
0.266
0.550
0.515
0.626
0.446
0.696
0.755
0.554
0.495
0.607
CoT
0.217
0.480
0.538
0.548
0.416
0.648
0.613
0.452
0.446
0.532
Multi-Step
0.250
0.482
0.452
0.486
0.394
0.503
0.603
0.530
0.425
0.500
Contriever
Vanilla
0.284
0.563
0.614
0.576
0.490
0.638
0.735
0.619
0.531
0.599
Appendix
Table 9: Evaluation of verbal confidence calibration performance (ECE and AUROC) on four datasets across varying retrievers and prompting strategies. Results show that the model consistently exhibits an average ECE greater than 0.4, indicating poor calibration performance.
HotpotQA
NQ
HotpotQA
NQ
Method
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Method
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Llama-3.1-8B-Instruct
Qwen2.5-7B-Instruct
Vanilla
0.358
0.630
0.389
0.655
Vanilla
0.333
0.692
0.363
0.707
CoT
0.340
0.672
0.370
0.679
CoT
0.327
0.712
0.349
0.685
Noise-aware
0.263
0.693
0.315
0.661
Noise-aware
0.285
0.673
0.316
0.650
Ensemble
0.316
0.648
0.364
0.683
Ensemble
0.318
0.718
0.363
0.707
Appendix
Table 10: Four passages OOD results. Lower ECE and higher AUROC indicate better calibration performance.
Setting
Pos
Bamboogle
HotpotQA
NQ
StrategyQA
Average
ECE
AUROC
ECE
AUROC
ECE
AUROC
ECE
AUROC
ECE
AUROC
gt_only
N/A
0.071
0.675
0.117
0.693
0.139
0.688
0.046
0.679
0.093
0.684
gt_with_noise/counterfactual
pos1
0.475
0.474
0.500
0.475
0.477
0.483
0.648
0.290
0.525
0.431
pos2
0.416
0.581
0.514
0.504
0.490
0.557
0.609
0.291
0.507
0.483
pos3
0.365
0.671
0.452
0.613
0.472
0.714
0.492
0.397
0.445
0.599
gt_with_noise/relevant
pos1
0.126
0.579
0.199
0.538
0.196
0.612
0.058
0.572
0.145
0.575
Appendix
Table 11: The table evaluate the impact of passage ordering on Llama-3.1-8B-Instruct ’s calibration performance. The “Setting” column defines the context structure, where gt_only refers to a noise-free baseline containing only the ground truth passage, while the gt_with_noise categories involve mixing the ground truth with specific types of noise (counterfactual, relevant, or irrelevant). The “Pos” column specifies the exact position (1st, 2nd, or 3rd) of the ground truth passage within the sequence of retrieved passages, designed to assess the model’s sensitivity to positional bias when processing mixed-quality contexts. The results indicate that calibration performance steadily declines as noise passages are added.
ECE
AUROC
Model
Baseline
NOVA
Δ
p
Sig.
Baseline
NOVA
Δ
p
Sig.
DeepSeek-R1-Distill-Llama-8B
0.407
0.311
+0.096
0.0001
0.657
0.679
+0.022
0.0425
**
DeepSeek-R1-Distill-Qwen-7B
0.437
0.344
+0.093
0.0001
0.654
0.723
+0.069
0.0001
Llama-3.1-8B-Instruct
0.411
0.266
+0.145
0.0001
0.605
0.751
+0.146
0.0001
Qwen2.5-7B-Instruct
0.366
0.264
+0.102
0.0001
0.730
0.768
+0.038
0.0052
Appendix
Table 12: Significance analysis comparing NOVA with Vanilla. Here, significance codes are defined as follows: *** p≤0.01 , ** p≤0.05 , * p≤0.10 , and ns indicates not significant.
ECE
AUROC
Model
Baseline
NOVA
Δ
p
Sig.
Baseline
NOVA
Δ
p
Sig.
DeepSeek-R1-Distill-Llama-8B
0.436
0.311
+0.125
0.0001
0.650
0.679
+0.029
0.0143
**
DeepSeek-R1-Distill-Qwen-7B
0.474
0.344
+0.130
0.0001
0.655
0.723
+0.068
0.0001
Llama-3.1-8B-Instruct
0.377
0.266
+0.111
0.0001
0.591
0.751
+0.160
0.0001
Qwen2.5-7B-Instruct
0.335
0.264
+0.071
0.0001
0.726
0.768
+0.042
0.0030
Appendix
Table 13: Significance analysis comparing NOVA with CoT. The significance codes are defined the same in Table 12
Setting
Pos
Bamboogle
HotpotQA
NQ
StrategyQA
Average
ECE
AUROC
ECE
AUROC
ECE
AUROC
ECE
AUROC
ECE
AUROC
gt_only
N/A
0.079
0.750
0.144
0.567
0.144
0.576
0.063
0.719
0.108
0.653
gt_with_noise/counterfactual
pos1
0.395
0.502
0.450
0.508
0.519
0.524
0.670
0.380
0.509
0.479
pos2
0.340
0.527
0.519
0.525
0.527
0.538
0.623
0.404
0.502
0.499
pos3
0.290
0.573
0.475
0.534
0.482
0.540
0.600
0.371
0.462
0.505
gt_with_noise/relevant
pos1
0.093
0.572
0.168
0.546
0.166
0.602
0.082
0.708
0.127
0.607
Appendix
Table 14: The table evaluate the impact of passage ordering on DeepSeek-R1-Distill-Llama-8B ’s calibration performance. The “Setting” column defines the context structure, where gt_only refers to a noise-free baseline containing only the ground truth passage, while the gt_with_noise categories involve mixing the ground truth with specific types of noise (counterfactual, relevant, or irrelevant). The “Pos” column specifies the exact position (1st, 2nd, or 3rd) of the ground truth passage within the sequence of retrieved passages, designed to assess the model’s sensitivity to positional bias when processing mixed-quality contexts. The results indicate that calibration performance steadily declines as noise passages are added.
Setting
bamboogle
hotpotqa
nq
strategyqa
Average
ECE
AUROC
ECE
AUROC
ECE
AUROC
ECE
AUROC
ECE
AUROC
noise_only / counterfactual
0.822
0.397
0.860
0.339
0.775
0.448
0.649
0.478
0.777
0.416
noise_only / irrelevant
0.225
0.851
0.227
0.772
0.293
0.776
0.345
0.554
0.273
0.738
noise_only / relevant
0.331
0.766
0.304
0.717
0.267
0.733
0.128
0.543
0.258
0.690
Appendix
Table 15: Llama3.1-8B performance in noise only settings.
Method
StrategyQA
HotpotQA
NQ
Bamboogle
Average
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Qwen2.5-7B-Instruct
Vanilla
0.398
0.689
0.391
0.712
0.438
0.710
0.236
0.809
0.366
0.730
A → NOVA
0.399
0.698
0.344
0.668
0.391
0.713
0.165
0.822
0.325
0.725
B → NOVA
0.371
0.702
0.376
0.742
0.417
0.732
0.219
0.855
0.346
0.758
C → NOVA
0.351
0.679
0.349
0.773
0.370
0.731
0.145
0.793
0.304
0.744
Appendix
Table 16: Comparison results across four datasets on Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct. Lower ECE and higher AUROC indicate better calibration and discrimination performance.
Method
StrategyQA
HotpotQA
NQ
Bamboogle
Average
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Llama-3.1-8B-Instruct
Vanilla
0.238
0.573
0.497
0.642
0.414
0.725
0.670
0.625
0.455
0.641
CoT
0.205
0.485
0.496
0.552
0.369
0.688
0.566
0.557
0.409
0.571
Noise-aware
0.229
0.546
0.329
0.671
0.360
0.679
0.495
0.680
0.353
0.644
Ensemble
0.130
0.551
0.391
0.665
0.376
0.720
0.515
0.616
0.353
0.638
Appendix
Table 17: Calibration performance of various models on four datasets with bm25-facts retriever. Scores in bold indicate the best performance, while underlined scores denote the second-best.
Model
UQ Method
StrategyQA
HotpotQA
NQ
Bamboogle
Average
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Llama-3.1-8B-Instruct
Base
Ensemble(3)
0.370
0.609
0.397
0.650
0.428
0.619
0.214
0.713
0.352
0.648
Self-freq
0.436
0.513
0.338
0.686
0.325
0.665
0.267
0.777
0.342
0.660
LexicSim
0.421
0.504
0.318
0.694
0.338
0.679
0.245
0.796
0.331
0.668
EigValLap
0.367
0.506
0.333
0.693
0.373
0.698
0.283
0.796
0.339
0.673
Appendix
Table 18: Performance evaluation of four specific sampling-based UQ methods (Ensemble, Self-freq, LexicSim, and Eig ValLap) on the Base and fine-tuned models. Bold indicates better performance across all tested post-hoc calibration strategies.
Method
StrategyQA
HotpotQA
NQ
Bamboogle
Average
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Llama-3.1-8B-Instruct
Vanilla
0.296
0.569
0.519
0.614
0.415
0.673
0.656
0.607
0.472
0.616
CoT
0.167
0.550
0.585
0.476
0.347
0.649
0.592
0.535
0.423
0.552
Noise-aware
0.218
0.566
0.261
0.711
0.314
0.652
0.478
0.693
0.318
0.655
Ensemble
0.110
0.595
0.416
0.631
0.364
0.633
0.525
0.620
0.354
0.620
Appendix
Table 19: Calibration performance of various models on four datasets with Contriever-facts retriever. Scores in bold indicate the best performance, while underlined scores denote the second-best.
Method
StrategyQA
HotpotQA
NQ
Bamboogle
Average
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
ECE ↓
AUROC ↑
Llama-3.1-8B-Instruct
Vanilla QA
0.200
0.524
0.640
0.636
0.515
0.666
0.796
0.523
0.538
0.587
RAG+Vanilla (BM25)
0.238
0.573
0.497
0.642
0.414
0.725
0.670
0.625
0.455
0.641
RAG+CoT (BM25)
0.205
0.485
0.496
0.552
0.369
0.688
0.566
0.557
0.409
0.571
RAG+Vanilla (Contriever)
0.296
0.569
0.519
0.614
0.415
0.673
0.656
0.607
0.472
0.616
Appendix
Table 20: ECE and AUROC comparison between Vanilla QA (no retrieval) and RAG settings (BM25 / Contriever retrievers, Vanilla / CoT prompts).
Figure 5: A comprehensive evaluation of the baselines and NOVA (R1: Vanilla, R2: CoT, R3: Noise-aware, R4: Label-only SFT, R5: NOVA ). (a) The spider plot illustrates the multi-dimensional performance across 5 evaluation criteria, highlighting the overall balance. (b) The bar plot provides a detailed mean and variance of the evaluation criteria breakdown.
Figure 6: Screenshot of an example query in our human annotation interface. For each query, annotators answer five questions. For each question, they are presented with five candidate responses and asked to rate each response on a scale from 1 to 5.
Figure 7: Case study setup illustrating a high-conflict retrieval scenario. The input consists of a query and three retrieved passages: the Ground Truth passage (Passage 1) is mixed with two Counterfactual passages (Passages 2 and 3) that support mutually exclusive incorrect answers (“Blargon-7” and “Omicron Persei 8”), testing the model’s ability to handle contradictory evidence.
Figure 8: Comparison of model responses under counterfactual noise. The Vanilla model (top) fails to resolve the conflict, hallucinating an incorrect answer with high confidence (80%). In contrast, NOVA (bottom) employs step-by-step reasoning to explicitly identify the contradictions among retrieved passages. By adhering to the Conflict Independence rule, it falls back to internal knowledge and assigns an appropriately low confidence score (10%), demonstrating superior calibration.
Figure 9: Prompt templates for the baseline methods. We employ three prompting strategies: Vanilla, Chain-of-Thought (CoT), and Multi-step. The specific instructions requiring step-by-step reasoning and step-level confidence estimation are highlighted in red. The placeholders {question} and {retrieved passages} represent the specific question and passages for one prompt.
Figure 10: The noise-aware prompt used in Table 2 . The placeholders {question} and {retrieved passages} represent the specific question and passages for one prompt.
Figure 11: The NOVA prompt used in Table 2 . The placeholders {question} and {retrieved passages} represent the specific question and passages for one prompt.
Figure 12: Noise generation prompt (for counterfactual noise ) used in Section § 5.1 . The placeholders {sentence_length} and {word_length} are calculated based on the length of the ground truth passage of each question, to make sure that our generated noise length is approximately the same with the ground truth passage. The concrete example represented by "[Example]" are presented in Figure 13 .
Figure 13: The Example partof the noise generation prompt for counterfactual noise . The placeholders {query} and {gt_answer} represent the specific question and answer pair to generate noise passages.
Figure 14: Noise generation prompt (for relevant noise ) used in Section § 5.1 . An example of this prompt is shown in Figure 15 . The placeholders {sentence_length} and {word_length} are calculated based on the length of the ground truth passage of each question, to make sure that our generated noise length is approximately the same with the ground truth passage. The placeholders {query} and {gt_answer} represent the specific question and answer pair to generate noise passages.
Figure 15: The Example part of the noise generation prompt for relevant noise . The placeholders {query} and {gt_answer} represent the specific question and answer pair to generate noise passages.
Figure 16: Noise generation prompt (for irrelevant noise ) used in Section § 5.1 . An example of this prompt is shown in Figure 17 . The placeholders {sentence_length} and {word_length} are calculated based on the length of the ground truth passage of each question, to make sure that our generated noise length is approximately the same with the ground truth passage. The placeholders {query} and {gt_answer} represent the specific question and answer pair to generate noise passages.
Figure 17: The Example part of the noise generation prompt for irrelevant noise . The placeholders {query}, {gt_answer} and {gt_passage} represent the specific question, answer and passage to generate noise passages.
Large language models (LLMs) often produce answers with high certainty even when they are incorrect, making reliable confidence estimation essential for deployment in real-world scenarios. Verbalized confidence, where models explicitly state their confidence in natural language, provides a flexible and user-facing uncertainty signal that can be applied even when token logits are unavailable. However, existing verbalized-confidence methods often optimize answer generation and confidence generation jointly, which can cause confidence-alignment objectives to interfere with answer accuracy. In this work, we propose a decoupled and order-aware framework for verbalized confidence calibration. Our method first generates an answer and then estimates confidence conditioned on the fixed question--answer pair, allowing confidence optimization without directly perturbing the answer-generation process. To align confidence with correctness likelihood, we construct a sampling-based surrogate from multiple model completions and optimize rank-based reinforcement learning objectives that encourage responses with higher estimated correctness likelihood to receive higher verbalized confidence. Experiments on reasoning and knowledge-intensive benchmarks show that our method improves calibration and failure prediction performance while largely preserving answer accuracy. These results demonstrate that verbalized confidence can be more reliably aligned by decoupling confidence estimation from answer generation and optimizing the relative ordering of confidence across responses.
Chen Li, Xiaoling Hu, Songzhu Zheng +2
Stony Brook University, NY, USA · Massachusetts General Hospital and Harvard Medical School, MA, USA · Morgan Stanley, NY, USA
Linguistic cues such as "I believe" and "probably" offer an intuitive interface for communicating confidence, yet a generalisable, principled calibration framework for linguistic confidence expressions remains underexplored. In particular, co-occurring linguistic cues, contextual variation, and subjective audience interpretation pose unique challenges. We therefore model linguistic confidence as a distribution over plausible perceived probability values that a statement is correct, capturing interpretation variability that scalar representations discard. Within this distributional framework, we introduce faithfulness as a complementary evaluation dimension and present Faithfulness Divergence (FD), an information-theoretic metric quantifying the surprise induced in audience beliefs upon truth revelation. Building on these foundations, we present Retrieval-Augmented Linguistic Calibration (RALC), a lightweight post-hoc pipeline that propagates calibrated confidence signals back into natural language via retrieval-augmented rewriting. Across three QA benchmarks and five LLM families, RALC improves in-domain faithfulness and calibration up to 66% and 58%, respectively, outperforming black-box and grey-box calibration baselines.
Yi-Fan Yeh, Linwei Tao, Minjing Dong +4
School of Computer Science University of Sydney Sydney, Australia · City University of Hong Kong Hong Kong · Shanghai Jiao Tong University Shanghai, China +1
Large language models (LLMs) can enhance factuality via retrieval-augmented generation (RAG), but applying RAG to every query is unnecessary when the model-only answer is reliable. This motivates cascaded RAG: each query is first handled by an LLM-only branch, escalated to a RAG fallback only if the primary branch is uncertain, and abstained from when neither branch is sufficiently trustworthy. However, calibrating such cascades stage by stage may be conservative, since the final utility depends on joint uncertainty thresholding of LLM-only and RAG. In this work, we develop BalanceRAG to certify threshold pairs at a target risk level. Given uncertainty scores from the two branches, BalanceRAG frames each threshold pair as an operating point on a two-dimensional lattice and identifies safe operating points using sequential graphical testing. This enables risk-adaptive threshold calibration, controlling the system-level error rate among accepted points, while retaining more examples. Furthermore, BalanceRAG extends to multi-risk calibration, allowing retrieval usage to be bounded together with the selection-conditioned risk. Experiments on three open-domain question answering (QA) benchmarks across multiple LLM backbones demonstrate that BalanceRAG meets prescribed risk levels, preserves higher coverage and more accepted correct examples, and reduces unnecessary retrieval calls compared with always-on RAG.
Zijun Jia, Yuanchang Ye, Sen Jia +6
Beihang University · Zhejiang University of Finance & Economics · Shenzhen Institute of Advanced Technology +1