Organizations: Oxford Dynamics Oxford · Ludwig-Maximilians-Universität München Munich · Institute for Artificial Intelligence, Data Analysis and Systems (AIDAS) School of Engineering Computing & Mathematics Oxford Brookes University, Oxford, UK
Large language models (LLMs) often produce fluent but incorrect answers with unwarranted confidence. A central limitation is that standard LLMs represent uncertainty through a single predictive distribution, conflating epistemic ignorance with genuine ambiguity. We introduce Credal Large Language Models (CLLMs): an ensemble of LoRA adapters induces a credal set whose lower and upper probabilities expose the spread of plausible predictive distributions rather than collapsing to a single softmax output. From this representation, we derive a single commitment rule: the model commits to an answer only when its lower probability exceeds the upper probability of every alternative, and otherwise returns the set of answers that no plausible predictor rules out. We apply this commitment rule at two depths: Credal Token Commitment (CTC) applies it to answer tokens from one ensemble forward pass, which decides constrained answers without any generation; for open-ended answers, credal decoding extends a partial answer only when no completed answer dominates it, so that the completions produced are those the plausible predictors license, and Credal Semantic Commitment (CSC) applies the rule to their meaning clusters. We evaluate CLLMs with Gemma-2-9B, Llama-3.1-8B and Qwen2.5-7B on OpenBookQA, CoQA, TriviaQA and ARC-Challenge. On multiple choice, CTC commits on 73-91% of questions at 89-98% accuracy, returns sets of 1.1-1.5 options containing the gold one on 89-98%, and its intervals contain the observed accuracy in 24 of 30 confidence bins without calibration; corrupted context lowers commitment from 87-92% to 65-71%, and on Gemma the credal bound detects corruption better than every baseline. On open-ended QA, CLLM outperforms semantic entropy and Laplace-LoRA at a fixed coverage by up to 19% and 9.5% absolute accuracy on CoQA and TriviaQA with context, for every backbone.
Figures & tables
Figure 1 : Left: one softmax over five candidate answers a - e is sharp whether the model knows the answer or merely settled on one. Right: the credal set P(A) ( Eq. 1 ) of five LoRA adapters gives each answer a lower (solid) and an upper (band) probability. Here the lower bound of a (dashed) is below the upper bound of b , so some adapter prefers b : the model should not commit to a alone.
Figure 2 : Credal Large Language Models. (1) M LoRA adapters on a frozen backbone give M next-token distributions. (2a) Plotted in the probability simplex, reduced for display to the three tokens that carry mass here, Paris , The and Lyon , they generate the credal set of Eq. 1 , their convex hull (shaded). The dashed region is the larger set of distributions consistent with the bounds alone; the two agree on the lower and upper probabilities of Eq. 2 , which is all the rule reads. One interval-dominance test, Eq. 3 , is then applied in both spaces. In the token space Paris and The both reach the largest lower bound, so CTC does not commit and returns the surviving set D1 . (2b) Credal decoding writes out the completions no adapter dominates ( Algorithm 1 ), the adapters themselves group them by meaning, and each meaning’s upper bound is widened by that adapter’s unexplored mass um ( Eq. 4 ) (dashed extension). (3) In the semantic space only c1 survives and its lower bound clears the largest unexplored mass, so CSC commits to the class c1={h1,h2} , that is to the answer Paris rather than to either sentence: the bounds are computed from the pooled mass of both completions. The two spaces are read independently on the same prompt. Here the token space does not close, because the adapters disagree about how the answer begins rather than about what it is, and the semantic space decides.
CoQA
TriviaQA (passage)
TriviaQA (closed book)
OpenBookQA
ARC-Challenge
Score
Source
Clean ↑
Corrupted ↑
Clean ↑
Corrupted ↑
Clean ↑
Corrupted ↑
Clean ↑
Corrupted ↑
Clean ↑
Corrupted ↑
Standard LLM (predictive entropy)
single model
0.786
0.739
0.850
0.568
0.729
0.495
0.866
0.538
0.762
0.472
Standard LLM (max prob)
single model
0.774
0.754
0.829
0.594
0.702
0.502
0.863
0.537
0.760
0.488
Semantic entropy (cosine)
single model, K=16
0.646
0.739
0.736
0.697
0.722
0.543
0.668
0.638
0.633
0.490
Semantic entropy (NLI)
single model, K=16
0.709
0.698
0.811
0.675
0.723
0.556
0.535
0.629
0.530
0.502
Bayesian-LoRA (KFAC)
posterior, S=20
0.812
0.827
0.845
0.694
0.792
0.597
0.931
0.642
0.819
0.512
Table 1 : Hallucination and corrupted-context detection (AUROC, Gemma-2-9B; corrupted columns for Llama-3.1-8B and Qwen2.5-7B in Fig. 5 ; 0.5 is chance and no score is inverted). Clean : does the score separate the questions the model’s own answer got wrong from the ones it got right ( Kuhn et al., 2023 ; Farquhar et al., 2024 ) ? Corrupted : does it separate corrupted inputs from intact ones? Where the answer rests on supplied evidence, the evidence is replaced (the passage in CoQA and TriviaQA, the supporting fact in OpenBookQA); otherwise the question is replaced and the answer side kept (closed-book TriviaQA and ARC-Challenge).
CoQA
TriviaQA (passage)
TriviaQA (closed book)
OpenBookQA
ARC-Challenge
Score
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Standard LLM (predictive entropy)
0.795
0.156
0.694
0.710
0.241
0.784
0.735
0.160
0.768
0.973
0.067
0.302
0.943
0.095
0.438
Standard LLM (max prob)
0.790
0.156
0.694
0.700
0.241
0.784
0.715
0.160
0.768
0.973
0.067
0.302
0.940
0.095
0.438
Semantic entropy (cosine)
0.780
0.198
3.141
0.720
0.287
3.610
0.755
0.180
2.480
0.948
0.070
1.237
0.915
0.107
2.043
Semantic entropy (NLI)
0.775
0.129
1.775
0.725
0.240
2.353
0.750
0.144
2.125
0.917
0.138
0.835
0.875
0.130
1.412
Bayesian-LoRA (KFAC)
0.835
0.271
0.645
0.900
0.435
0.854
0.755
0.110
0.537
0.993
0.009
0.130
0.940
0.068
0.322
Table 2 : Selective accuracy and calibration (Gemma-2-9B; Llama-3.1-8B and Qwen2.5-7B in Tabs. 5 and 6 ). Acc.@80%: accuracy at 80% coverage when questions are ranked by the score. ECE and NLL: calibration of each method’s own per-answer confidence; for CLLM this is the midpoint of the interval, since the lower bound is a guaranteed floor and under-confident by construction. Methods built on the same adapters select the same answer and differ only in the score. ARC-Challenge uses the OpenBookQA-trained adapters.
Figure 3 : The rule decides where a tuned threshold would, and withdraws when the evidence is corrupted (Gemma-2-9B on CoQA; Llama-3.1-8B and Qwen2.5-7B, and TriviaQA, in Fig. 6 and Fig. 7 ). Left , risk-coverage: solid is clean prompts, dashed the same prompts with the passage replaced; the circle and square mark the rule’s own operating point, chosen without a threshold. Corrupting the passage moves it from 47.6% coverage to 3.6% , while a ranker’s curve only shifts down. Right , each method thresholded on clean prompts to the coverage the rule chose: bars are what it still answers. Semantic entropy is zero on most prompts, so no threshold reaches that coverage and its lowest attainable one is shown (hatched).
Figure 4 : CTC withdraws under corrupted evidence rather than degrading . Each marker is one backbone at its commit rate and its accuracy when committed. Filled: OpenBookQA, clean; hollow: the same questions with the passage corrupted; blue: ARC-Challenge, answered with the OpenBookQA-trained adapters.
Backbone
Committed on harmful
Committed on benign
P(a∗) harmful / benign
Standard LLM entropy
Llama-3.1-8B
71.6%
30.4%
0.49 / 0.28
0.03 / 0.91
Gemma-2-9B
97.2%
38.4%
0.93 / 0.32
0.02 / 0.69
Qwen2.5-7B
33.2%
22.8%
0.25 / 0.19
0.02 / 0.33
Table 3 : Harmful prompts ( 250 AdvBench harmful against 250 benign Dolly prompts, CoQA-trained ensembles). Committed answers on harmful prompts are refusals. The last two columns are means on each half: the ensemble is more confident and the single model less uncertain on harmful prompts, so neither can flag harm.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Token
f1
f2
f3
f4
f5
P
P
In D1
special
0.684
0.029
0.245
0.925
0.938
0.029
0.938
yes
Special
0.222
0.847
0.666
0.046
0.013
0.013
0.847
yes
food
0.044
0.005
0.016
0.022
0.036
0.005
0.044
yes
They
0.001
0.037
0.002
0.000
0.000
0.000
0.037
yes
fifth token
0.018
0.003
0.002
0.004
0.006
0.002
0.018
no
sixth token
0.004
0.014
0.002
0.000
0.000
0.000
0.014
no
Appendix
Table 4: The token space of the worked example. Probability of each of the six most likely first answer tokens under each of the five adapters, with the bounds of Eq. 2 . The largest lower probability is 0.029 , so every token whose upper probability reaches 0.029 survives the rule of Eq. 3 . Four tokens survive, so the token space does not commit.
CoQA
TriviaQA (passage)
TriviaQA (closed book)
OpenBookQA
ARC-Challenge
Score
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Standard LLM (predictive entropy)
0.775
0.083
0.511
0.895
0.075
0.342
0.785
0.117
0.563
0.965
0.040
0.240
0.858
0.104
0.504
Standard LLM (max prob)
0.770
0.083
0.511
0.890
0.075
0.342
0.785
0.117
0.563
0.968
0.040
0.240
0.860
0.104
0.504
Semantic entropy (cosine)
0.755
0.154
1.825
0.885
0.070
0.765
0.795
0.136
1.127
0.968
0.047
0.622
0.860
0.131
1.458
Semantic entropy (NLI)
0.745
0.177
0.871
0.905
0.291
0.565
0.810
0.229
0.829
0.945
0.175
0.447
0.860
0.041
0.821
Bayesian-LoRA (KFAC)
0.745
0.062
0.510
0.890
0.438
0.855
0.785
0.103
0.543
0.963
0.025
0.224
0.855
0.073
0.415
Appendix
Table 5 : Comparison with baselines on every dataset, Llama-3.1-8B (CoQA, TriviaQA with and without the passage, OpenBookQA and ARC-Challenge). Each dataset carries three measures: Acc.@80% , the accuracy of the selected answer when the questions are ranked by the method’s score and the easiest 80% are answered, higher is better; ECE and NLL , the expected calibration error and negative log likelihood of the method’s own per-answer confidence, lower is better in both. For the credal row the confidence is the midpoint of the interval. Rows that select the same answer share that answer’s confidence, so the two Standard LLM rows and the two LoRA Ensemble rows carry the same calibration. CLLM is one row, at the depth the answer type fixes. Bold = best per column, underlined = second-best; the grey row is ours.
CoQA
TriviaQA (passage)
TriviaQA (closed book)
OpenBookQA
ARC-Challenge
Score
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Acc.@80% ↑
ECE ↓
NLL ↓
Standard LLM (predictive entropy)
0.740
0.212
1.173
0.735
0.225
1.021
0.545
0.370
2.418
0.975
0.090
0.940
0.938
0.109
1.459
Standard LLM (max prob)
0.740
0.212
1.173
0.715
0.225
1.021
0.530
0.370
2.418
0.975
0.090
0.940
0.938
0.109
1.459
Semantic entropy (cosine)
0.745
0.207
4.255
0.750
0.208
3.084
0.585
0.371
6.550
0.920
0.091
1.752
0.897
0.113
2.341
Semantic entropy (NLI)
0.770
0.145
2.213
0.790
0.100
1.252
0.620
0.130
1.003
0.900
0.138
1.645
0.885
0.131
1.331
Bayesian-LoRA (KFAC)
0.770
0.128
0.665
0.845
0.067
0.485
0.555
0.225
1.119
0.975
0.085
0.831
0.935
0.106
1.392
Appendix
Table 6 : Comparison with baselines on every dataset, Qwen2.5-7B (CoQA, TriviaQA with and without the passage, OpenBookQA and ARC-Challenge). Each dataset carries three measures: Acc.@80% , the accuracy of the selected answer when the questions are ranked by the method’s score and the easiest 80% are answered, higher is better; ECE and NLL , the expected calibration error and negative log likelihood of the method’s own per-answer confidence, lower is better in both. For the credal row the confidence is the midpoint of the interval. Rows that select the same answer share that answer’s confidence, so the two Standard LLM rows and the two LoRA Ensemble rows carry the same calibration. CLLM is one row, at the depth the answer type fixes. Bold = best per column, underlined = second-best; the grey row is ours.
Dataset
Backbone
Standard LLM (predictive entropy) ↑
LoRA Ensemble (predictive entropy) ↑
CLLM ↑
OpenBookQA
Llama-3.1-8B
0.965
0.965
0.980
OpenBookQA
Gemma-2-9B
0.973
0.995
0.998
OpenBookQA
Qwen2.5-7B
0.975
0.978
0.980
ARC-Challenge
Llama-3.1-8B
0.858
0.810
0.890
ARC-Challenge
Gemma-2-9B
0.943
0.950
0.953
ARC-Challenge
Qwen2.5-7B
0.938
0.940
0.938
Appendix
Table 7: Selective accuracy at 80% coverage on the four-option setting, OpenBookQA and ARC-Challenge ( N=500 ; ensemble scores rank the representative answer, the Standard LLM ranks its greedy answer).
Figure 5 : Corrupted-context detection across all baseline families , AUROC on 250 clean against 250 corrupted prompts per cell. Credal quantities are shown against the mean-pooled ensemble scores, the two posterior approximations and the single-model scores, for both open-ended datasets and all three backbones.
CoQA
TriviaQA (passage)
Statistic (AUROC ↑ )
Llama
Gemma
Qwen
Llama
Gemma
Qwen
Single model, predictive entropy
0.669
0.739
0.557
0.717
0.568
0.528
Mean-pooled ensemble, predictive entropy
0.837
0.830
0.797
0.817
0.683
0.749
Mean-pooled ensemble, mutual information
0.863
0.854
0.786
0.824
0.723
0.836
Credal set, entropy of the interpolating distribution
0.876
0.860
0.809
0.813
0.688
0.674
Credal set, upper entropy
0.855
0.854
0.805
0.779
0.690
0.736
Appendix
Table 8 : Scalar statistics read from the token-space credal set, as corrupted-versus-clean detectors (AUROC; 250 clean and 250 corrupted prompts per cell, M=5 ). Every ensemble and credal row is read from the same five adapters. The rows that also appear in Tab. 1 carry the same values there. Bold = best per column, underlined = second best; grey rows are ours.
Figure 6 : Risk-coverage curves on clean prompts , CoQA and TriviaQA with the passage, all three backbones. The horizontal axis is the fraction of questions answered when they are ranked by each method’s score, the vertical axis the accuracy of the answers given; a curve that stays high as it moves right ranks its own errors well. The filled marker is the operating point our rule chose for itself, with no threshold; the hollow marker is the LoRA Ensemble at the same coverage.
Figure 7 : Abstention at matched clean coverage , both open-ended datasets, all three backbones. Each method is thresholded on clean prompts to the coverage our rule chose for itself, then shown the corrupted prompts; the bars are the fraction of questions it still answers. A method that answers as much of the corrupted split as of the clean one has not noticed the intervention.
Matched size
Fixed size
Dataset
Backbone
Our mean size ↓
Ours ↑
Top- k at our size ↑
k=1
k=2
k=3
CoQA
Llama-3.1-8B
2.52
87
87
75
83
84
CoQA
Gemma-2-9B
2.65
90
90
79
87
89
CoQA
Qwen2.5-7B
5.94
89
89
76
82
85
TriviaQA (passage)
Llama-3.1-8B
3.78
91
91
85
88
90
TriviaQA (passage)
Gemma-2-9B
2.51
88
88
80
85
87
Appendix
Table 9 : The size, not the ranking (clean prompts, gold-in-set, per cent). Given our per-question size the ensemble’s set is ours on almost every prompt, so those columns agree; against any k fixed in advance the credal set is ahead or level, on CoQA while showing fewer meanings.
Dataset
Backbone
Commit rate
Mean ∣D1∣↓
Gold in set ↑
Acc. at the rule’s coverage
CTC (ours)
LoRA Ensemble (pred. entropy)
OpenBookQA
Llama-3.1-8B
0.830
1.29
0.966
0.971
0.959
OpenBookQA
Gemma-2-9B
0.864
1.30
0.984
0.984
0.986
OpenBookQA
Qwen2.5-7B
0.906
1.17
0.956
0.954
0.954
ARC-Challenge
Llama-3.1-8B
0.734
1.50
0.888
0.905
0.888
ARC-Challenge
Gemma-2-9B
0.882
1.20
0.930
0.934
0.927
Appendix
Table 10 : Credal Token Commitment on the four-option setting , the commit rates behind Fig. 4 (OpenBookQA and ARC-Challenge, N=500 , M=5 ; ARC-Challenge uses the OpenBookQA-trained adapters). Commit =∣D1∣=1 . The final block is the paired OpenBookQA split, the same 500 questions with and without a corrupted passage, so its clean rows differ slightly from the first block, which uses the standard test split.
Figure 8 : Interval calibration on the four-option setting. Questions are binned by the midpoint of the selected answer’s interval; the bar is the bin’s mean [P,P] , the dot the observed accuracy (red when outside the interval), the number the bin count; bins with fewer than 10 questions are dropped. The table form is Tab. 11 .
All populated bins
Top bin
Dataset
Backbone
Bins
Inside ↑
Below P↓
Mean width ↓
Share of questions
Accuracy
Mean [P,P]
OpenBookQA
Llama-3.1-8B
6
6
0
0.447
0.71
0.99
[0.96,1.00]
OpenBookQA
Gemma-2-9B
4
4
0
0.462
0.86
0.99
[0.98,1.00]
OpenBookQA
Qwen2.5-7B
4
3
1
0.423
0.88
0.97
[0.98,1.00]
ARC-Challenge
Llama-3.1-8B
6
5
1
0.399
0.58
0.95
[0.96,1.00]
ARC-Challenge
Gemma-2-9B
5
3
2
0.468
0.83
0.95
[0.97,1.00]
Appendix
Table 11: Interval calibration on the four-option setting , OpenBookQA and ARC-Challenge. All questions binned by the midpoint of the representative answer’s interval ( 10 bins; bins with fewer than 10 questions dropped). A bin is “inside” when its accuracy lies within the bin’s mean [P(a∗),P(a∗)] . The last three columns describe the top bin.
Setting
M
Commit
Acc ∣ commit ↑
Gold in set ↑
Set size ↓
Acc (rep.) ↑
Unexpl.
Llama-3.1-8B
CoQA, clean
5
40.4
95
87
2.52
75
0.51
CoQA, corrupted context
5
3.6
22
24
3.59
14
0.88
TriviaQA, clean
5
1.2
100
91
3.78
86
0.80
TriviaQA, corrupted context
5
0.0
–
76
4.47
67
0.96
TriviaQA closed-book
5
32.0
96
88
2.80
77
0.58
Appendix
Table 12 : Credal Semantic Commitment on CoQA and TriviaQA, clean and with the passage corrupted, and TriviaQA closed-book (fine partition, unexplored mass carried; M=5 , 250 prompts per cell). Percentages except set size (visible meanings) and Unexpl. (mean over adapters of the mass the search did not reach); gold in set: the presented set contains the gold meaning; Acc (rep.): accuracy of the representative meaning ( argmaxP ).
Dataset
Backbone
Commit rate, clean
Commit rate, corrupted
Median ∣D1∣ clean / corrupted
AUROC of ∣D1∣
OpenBookQA
Gemma-2-9B
0.898
0.706
1 / 1
0.881
OpenBookQA
Llama-3.1-8B
0.868
0.652
1 / 1
0.732
OpenBookQA
Qwen2.5-7B
0.916
0.670
1 / 1
0.805
CoQA
Gemma-2-9B
0.236
0.044
2 / 13
0.830
CoQA
Llama-3.1-8B
0.304
0.052
2 / 10
0.832
CoQA
Qwen2.5-7B
0.204
0.036
3 / 20
0.777
Appendix
Table 13 : Token-space commitment under corrupted context ( 250 clean and 250 corrupted prompts per CoQA and TriviaQA cell, and the 500 paired OpenBookQA questions of Tab. 10 ; space = whole vocabulary, top- 100 tokens by upper bound plus the remaining mass). A prompt is corrupted when its evidential passage is replaced or perturbed. OpenBookQA rows use the first-token caches of Tab. 10 ; CoQA and TriviaQA rows use the credal-decoding caches of Tab. 12 ( M=5 ). The last column is the corrupted-versus-clean AUROC of the set size.
Setting
Both commit
Token space only
Semantic space only
Neither
Llama-3.1-8B
CoQA, clean
20
10
20
50
CoQA, corrupted context
3
2
0
94
TriviaQA, clean
0
0
1
99
TriviaQA, corrupted context
0
0
0
100
TriviaQA closed-book
24
15
8
53
Appendix
Table 14: Agreement between depths on CoQA and TriviaQA, with the passage and closed-book (percent of questions): both CTC and CSC commit; only the token space; only the semantic space; neither.
Figure 9 : Agreement between the two depths, per cell (the data of Tab. 14 ).
CoQA
TriviaQA (passage)
TriviaQA (closed book)
OpenBookQA
ARC-Challenge
Depth (AUROC ↑ )
Candidates
Llama
Gemma
Qwen2.5
Llama
Gemma
Qwen2.5
Llama
Gemma
Qwen2.5
Llama
Gemma
Qwen2.5
Llama
Gemma
Qwen2.5
CLLM (CTC)
token space, first answer token
0.754
0.752
0.714
0.711
0.744
0.755
0.764
0.706
0.536
0.917
0.942
0.917
0.781
0.830
0.807
CLLM (CSC)
semantic space, meanings
0.830
0.844
0.740
0.822
0.850
0.867
0.833
0.833
0.783
0.917
0.942
0.917
0.781
0.830
0.807
Appendix
Table 15 : The same rule at both depths, on every dataset (AUROC of each depth’s score against the model’s own answer being wrong, higher is better; CoQA, TriviaQA with and without the passage, OpenBookQA and ARC-Challenge). The main tables report one CLLM row, whose depth is fixed by the answer type rather than chosen per cell: the semantic space on free-form answers, the token space on the two four-option datasets, where a one-letter answer is its own meaning and the two depths are the same quantity. This table is the check on that rule. On free-form answers the semantic space is ahead in all nine cells, by 3 to 25 points. Bold = better of the two depths per column; grey rows are ours.
M
CTC commit
CSC commit
Acc ∣ commit
Gold in set
Set size
Unexpl.
2
38.6
43.4
90
84
1.79
0.50
3
31.4
41.6
91
88
2.08
0.50
4
28.2
40.6
93
90
2.33
0.51
5
26.6
40.8
92
89
2.34
0.51
Appendix
Table 16: Number of adapters (CoQA development split, 500 prompts, Llama-3.1-8B, sequence-level expansion, fine partition). Percentages except set size (visible meanings) and unexplored mass.
Figure 10 : Number of adapters and expansion rule (CoQA development split, 500 prompts, nested subsets M=2,3,4,5 ; hollow marker: token-level expansion at M=5 ). Token-space commitment falls with M on all three backbones while gold-in-set rises or holds.
Backbone
Expansion
CSC commit
Acc. when committed
Gold in set
Steps/prompt
Completed/prompt
Llama-3.1-8B
decision
40.8
92.2
89.2
72.4
10.9
Llama-3.1-8B
token
35.0
92.6
87.4
83.7
14.9
Gemma-2-9B
decision
48.8
92.6
88.6
73.0
23.9
Gemma-2-9B
token
43.2
93.1
87.2
65.1
27.2
Qwen2.5-7B
decision
29.0
93.8
88.8
97.5
14.9
Qwen2.5-7B
token
28.0
94.3
88.0
106.2
17.4
Appendix
Table 17 : Expansion rule (CoQA development split, 500 prompts, M=5 , B=8 ). Decision : prune a partial sequence when its upper mass falls below the best completed lower mass. Token : keep only undominated next tokens. Commitment and gold-in-set improve everywhere; the step count falls on two backbones of three.
Meanings per prompt
Impurity
Dataset
Backbone
sure
possible
sure
possible
CoQA
Llama-3.1-8B
2.98
1.41
30%
62%
CoQA
Gemma-2-9B
3.32
1.42
30%
64%
CoQA
Qwen2.5-7B
6.88
3.94
19%
40%
TriviaQA
Llama-3.1-8B
3.80
1.01
60%
88%
TriviaQA
Gemma-2-9B
2.52
1.00
38%
66%
Appendix
Table 18 : Sure against possible meaning partitions on clean test prompts ( 250 per cell). Meanings: the mean number of meanings per prompt. Impurity: the percentage of prompts on which one meaning contains both a gold-matching and a non-matching candidate. Every result in the paper uses the sure partition.
Rule as stated
κ=0.5
κ=1
κ=2
κ=4
Dataset
Backbone
Cov.
Acc. ↑
Cov.
Acc. ↑
Cov.
Acc. ↑
Cov.
Acc. ↑
Cov.
Acc. ↑
CoQA
Llama-3.1-8B
0.40
0.95
0.54
0.93
0.40
0.95
0.28
0.97
0.16
1.00
CoQA
Gemma-2-9B
0.48
0.96
0.56
0.94
0.47
0.96
0.31
0.99
0.18
1.00
CoQA
Qwen2.5-7B
0.31
0.91
0.37
0.90
0.31
0.91
0.22
0.98
0.14
1.00
TriviaQA
Llama-3.1-8B
0.01
1.00
0.07
1.00
0.01
1.00
0.00
–
0.00
–
TriviaQA
Gemma-2-9B
0.01
1.00
0.17
0.98
0.01
1.00
0.00
–
0.00
–
Appendix
Table 19 : Cost-sensitive commitment on clean CoQA and TriviaQA prompts ( 250 per cell). Cov.: fraction of questions answered; Acc.: accuracy of the committed answer. “Rule as stated” is the parameter-free rule of Eq. 3 with the unexplored-mass condition; the remaining columns commit when P(c∗)≥κ/(1+κ) . A dash marks a column that commits on no question in that cell.
LLMs' overconfidence, particularly when hallucinating, poses a significant challenge for the deployment of the models in safety-critical settings and makes a reliable estimation of uncertainty necessary. Existing approaches for uncertainty quantification typically prioritize lexical or probabilistic measures; however, these techniques often ignore the semantic variance of different responses with similar meaning. In this paper, we propose Adaptive Conformal Semantic Entropy (ACSE), a method for estimating prompt-level uncertainty by adaptively measuring semantic dispersion in LLMs outputs. Our uncertainty scoring function is based on clustering semantic entropy of multiple diverse responses to the same prompt. The function adaptively adjusts the uncertainty score based on semantic features of each cluster. To ensure statistical reliability of our score, we use conformal calibration to apply a decision rule to accept/abstain the prompts, providing a finite-sample, distribution-free guarantee such that the error rate among the accepted responses remains bounded by a user-specified tolerance. Our extensive experimental evaluations using different LLMs and datasets, demonstrate that our approach consistently outperforms state-of-the-art uncertainty quantification baselines using discriminative performance, conformal guarantees, and probabilistic calibration indicators. As a highlight, for TriviaQA dataset, AUROC of our approach is 0.88 compared to 0.65 produced by the token entropy approach.
Hamed Karimi, Vaishali Meyappan, Reza Samavi
Toronto Metropolitan University, Toronto, Ontario, Canada · Vector Institute, Toronto, Ontario, Canada
Large language models (LLMs) are increasingly deployed in question answering (QA) systems, yet they may generate hallucinated or misaligned responses without reliable confidence estimates. Uncertainty quantification (UQ) offers a natural basis for selective answering, where a system answers only when its prediction is deemed reliable and abstains otherwise. However, existing uncertainty scores for LLMs are often heuristic: a threshold chosen on such scores does not, by itself, provide statistical guarantees on the error rate among accepted answers. We propose CIC, a confidence-interval-based calibration framework that converts arbitrary uncertainty scores into risk-controlled selective answering rules. Given a held-out calibration set, CIC evaluates each generated response using an application-specific alignment criterion and associates it with an uncertainty score and a binary error label. For each candidate uncertainty threshold, CIC estimates the acceptance-conditioned error rate and constructs a high-probability upper confidence bound using either Hoeffding-style or Clopper-Pearson confidence intervals. It then selects the largest threshold whose upper bound is below a user-specified risk level α, thereby maximizing the answering rate subject to a finite-sample reliability constraint. Under exchangeability, CIC guarantees with probability at least 1−δ that the selected threshold, if non-null, controls the error rate among accepted answers at level α. We evaluate CIC on both closed-ended and open-ended QA benchmarks across seven LLMs and multiple uncertainty estimators. Experimental results show that CIC consistently achieves valid risk control while retaining strong answering efficiency, providing a practical and statistically grounded mechanism for deploying LLMs in reliability-sensitive QA workflows.
Large language models can produce fluent answers when their factual support is weak. This paper introduces Chain-of-Self-Questioning (CoSQ), a prompt-only framework that makes answer commitment conditional on an explicit assessment of the information required to answer a question. We evaluate three CoSQ variants under seventeen conditions on the 817-item TruthfulQA multiple-choice validation set using eleven open-weight and hosted model families. In the final balanced-option protocol, Grounded-CoSQ at τ=0.90 reduces the mean unconditional wrong-commitment rate from 13.1% under chain-of-thought prompting to 8.9%, a 32.1% relative reduction, while increasing answered accuracy from 86.9% to 89.7% and answering 87.6% of questions. Both improvements hold for all eleven models and at every evaluated threshold. Critical-CoSQ and Adaptive-CoSQ provide neighboring operating points with 88.6% and 86.5% coverage, respectively, while remaining more reliable than the baseline. A secondary Natural Questions Short-Answer evaluation provides convergent open-form evidence. These findings show that self-assessment can support explicit, tunable answer-or-abstain decisions when an unsupported commitment is more costly than referral or review.
Ali Şenol
Department of Computer Engineering, Tarsus University, Tarsus, Türkiye