Large language models (LLMs) increasingly mediate how people access and reason with information, yet factual reliability is usually evaluated one judgment at a time. We introduce graded belief stability, a relational measure of how well a belief persists within an LLM's broader belief system. Unlike individual belief probability, it asks whether support for a claim persists when that claim is considered alongside the model's other epistemic commitments. We operationalize this idea with a Direct Conditional estimator that uses internal model representations to estimate conditional belief probabilities. Across 12 LLMs and three domains, lower-stability beliefs exhibit greater mean behavioral movement under conversational challenge in 83.3% of model-domain settings after matching on individual belief probability. Graded belief stability therefore extends reliability assessment beyond how strongly an LLM supports a claim to how robustly that belief is supported within its broader system of beliefs.
Figures & tables
Official Name
Short Name
# Layers
# Parameters
Release Date
Source
Citation
Gemma- 7 b-it
gemma-7b
28
8.54 B
Feb 21 , 2024
Google
[ 22 ]
Gemma- 2 - 9 b-it
gemma-2-9b
42
9.24 B
Jun 27 , 2024
Google
[ 22 ]
Gemma- 2 - 27 b-it
gemma-2-27b
46
27.23 B
Jun 27 , 2024
Google
[ 22 ]
Llama- 3.2 - 3 b-Instruct
llama-3.2-3b
28
3.21 B
Sep 25 , 2024
Meta
[ 23 ]
Llama- 3.1 - 8 b-Instruct
llama-3.1-8b
32
8.03 B
Jul 23 , 2024
Meta
[ 23 ]
Llama- 3.1 - 70 b-Instruct
llama-3.1-70b
80
70.55 B
Jul 23 , 2024
Meta
[ 23 ]
Table 1 : LLMs used in graded stability experiments. We list the official names of the LLMs according to the HuggingFace repository [ 45 ] , where they are publicly available. We further specify the shortened name used throughout the paper, the number of layers, parameter count, release date, source organization, and official citation.
Dataset
True
False
Synthetic
Examples
City Locations
A: 1392 N: 1376
A: 1358 N: 1374
A: 876 N: 876
T. The city of Surat is located in India. F. The city of Palembang is located in the Dominican Republic. S. The city of Norminsk is located in Jamoates.
Medical Indications
A: 1439 N: 1522
A: 1523 N: 1419
A: 478 N: 522
T. Pentobarbital is indicated for the treatment of insomnia. F. Vancomycin is not indicated for the treatment of lower respiratory tract infections. S. Alumil is indicated for the treatment of reticers.
Word Definitions
A: 1234 N: 1235
A: 1277 N: 1254
A: 1747 N: 1753
T. Hoagy is a synonym of an Italian sandwich. F. Decalogue is an astronomer. S. Dostab is a scencer.
Table 2 : Summary of datasets and statement types. Number of affirmative (A) and negated (N) statements across the three domains, along with examples. Each dataset includes True (T), False (F), and Synthetic (S) statements. Synthetic statements serve as Neither statements constructed to minimize prior LLM exposure. An earlier version of this table was introduced in [ 38 ] .
si
sj
si→sj
T
T
T
T
N
N
T
F
F
N
T
T
N
N
N
N
F
F
Table 3 : Cooper–Cantwell conditional truth table. We report whether the conditional si→sj≡sj∣si is True (T), False (F), or Neither (N) based on the truth values of si and sj according to the Cooper–Cantwell conditional [ 11 , 8 ] . When the antecedent si is True or Neither , the conditional takes the truth value of the consequent sj .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Description
M
A fixed large language model (LLM).
s
Natural-language statement.
P
A proposition believed by M whose stability is evaluated.
x
A non-disbelieved proposition used as conditioning information.
c∈{T,F,N}
Trivalent veracity class: True ( T ), False ( F ), or Neither ( N ).
K
Number of probe classes.
Appendix
Table A1 : Notation. Summary of the principal mathematical symbols used throughout the manuscript.
Dataset
Train
Calibration
Test
Total
City Locations
3999(0.55)
1398(0.19)
1855(0.26)
7252(1.00)
Medical Indications
3849(0.56)
1327(0.19)
1727(0.25)
6903(1.00)
Word Definitions
4717(0.55)
1628(0.19)
2155(0.25)
8500(1.00)
Appendix
Table A2 : Dataset splits. Number of statements used for training, calibration, and testing. Proportions of the full dataset are reported in parentheses. A version of this table appears in [ 38 ] .
Official Name
Short Name
# Layers
# Parameters
Release Date
Source
Citation
Gemma- 7 b
gemma-7b (b)
28
8.54 B
Feb 21 , 2024
Google
[ 22 ]
Gemma- 2 - 9 b
gemma-2-9b (b)
42
9.24 B
Jun 27 , 2024
Google
[ 22 ]
Gemma- 2 - 27 b
gemma-2-27b (b)
46
27.23 B
Jun 27 , 2024
Google
[ 22 ]
Llama- 3.2 - 3 b
llama-3.2-3b (b)
28
3.21 B
Sep 25 , 2024
Meta
[ 23 ]
Llama- 3.1 - 8 b
llama-3.1-8b (b)
32
8.03 B
Jul 23 , 2024
Meta
[ 23 ]
Llama- 3.1 - 70 b
llama-3.1-70b (b)
80
70.55 B
Jul 23 , 2024
Meta
[ 23 ]
Appendix
Table A3 : Base LLMs used in graded stability experiments. We list the official names of the LLMs according to the HuggingFace repository [ 45 ] , where they are publicly available. We further specify the shortened name used throughout the paper, the number of layers, parameter count, release date, source organization, and official citation.
Model
City Locations
Medical Indications
Word Definitions
llama-3.2-3b
9
11
13
llama-3.1-8b
12
19
12
llama-3.1-70b
34
34
17
gemma-7b
19
16
17
gemma-2-9b
23
20
21
gemma-2-27b
15
19
18
Appendix
Table A4 : Selected sAwMIL layers. Zero-indexed layers selected independently for each instruction-tuned LLM and dataset by minimizing three-class calibration log loss. The selected layer is subsequently used for the corresponding individual-statement, Direct Conditional, and Joint sAwMIL probes.
Belief sets
Conditional-pair coverage
Dataset
Model
∣BM∣
∣XM∣
tM
Candidate
Direct
Joint
Joint excl.
City Locations
llama-3.2-3b
632
1,217
0.506
768,512
768,512
768,512
0(0.0000%)
llama-3.1-8b
629
1,213
0.543
762,348
762,348
762,348
0(0.0000%)
llama-3.1-70b
628
1,212
0.478
760,508
760,508
760,471
37(0.0049%)
gemma-7b
615
1,199
0.513
736,770
736,770
736,770
0(0.0000%)
gemma-2-9b
626
1,211
0.537
757,460
757,460
756,727
733(0.0968%)
Appendix
Table A5: Belief-set and conditional-pair coverage using sAwMIL . For each LLM and domain, ∣BM∣ and ∣XM∣ denote the numbers of beliefs and non-disbeliefs, respectively, and tM is the empirical belief threshold. Candidate is the number ∣BM∣(∣XM∣−1) of constructed (P,x) pairs after excluding self-pairs. Direct and Joint report the numbers of valid conditional-probability estimates. Joint excl. reports the number and percentage of candidate pairs excluded because the Joint-to-Conditional denominator is at most 10−12 .
Dataset
Model
Exact
πTdirect(P,x)>πT(P)
πFdirect(P,x)>πF(P)
πF(x)>πNdirect(P,x)
πNdirect(P,x)>πN(P)+πF(x)
City Locations
llama-3.2-3b
0.000%
2.7%(0.095)
90.4%(0.154)
3.4%(0.122)
96.6%(0.207)
llama-3.1-8b
0.000%
5.2%(0.054)
88.0%(0.051)
5.8%(0.047)
94.2%(0.080)
llama-3.1-70b
<0.001%
7.9%(0.031)
80.5%(0.035)
11.4%(0.026)
88.5%(0.047)
gemma-7b
<0.001%
3.8%(0.107)
92.4%(0.159)
7.9%(0.091)
92.1%(0.102)
gemma-2-9b
0.000%
3.1%(0.050)
93.9%(0.087)
11.0%(0.036)
89.0%(0.046)
gemma-2-27b
0.000%
2.4%(0.067)
93.7%(0.104)
5.9%(0.073)
94.1%(0.109)
Appendix
Table A6: Exact CCK coherence and constraint violations for Direct Conditional estimates using sAwMIL . For each LLM and domain, we report the percentage of (P,x) pairs satisfying all four CCK constraints simultaneously within tolerance 10−8 . Here, 0.000% signifies exactly zero fully coherent pairs, while <0.001% means a small number of exactly coherent pairs exists. Each remaining column reports the percentage of pairs violating each constraint, with the mean excess across violating pairs shown in parentheses. Exact coherence is rare across all models and domains, with no model–domain combination exceeding 0.206% . Violations occur most frequently for the constraints on conditional False and Neither probability mass.
Dataset
Model
n
∣ΔπT∣
∣Δγ∣
City Locations
llama-3.2-3b
315
0.00435(0.01055)
0.235(0.173)
llama-3.1-8b
305
0.00181(0.00556)
0.090(0.122)
llama-3.1-70b
287
0.00483(0.02893)
0.066(0.116)
gemma-7b
304
0.00377(0.01042)
0.175(0.175)
gemma-2-9b
291
0.00370(0.01686)
0.082(0.141)
gemma-2-27b
316
0.00362(0.01372)
0.133(0.143)
Appendix
Table A7: Behavioral matching quality. For each LLM and domain, n is the number of matched proposition pairs, ∣ΔπT∣ is the absolute difference in belief probability between the two propositions in each pair, and ∣Δγ∣ is their absolute difference in graded stability. Values for ∣ΔπT∣ and ∣Δγ∣ report the mean with standard deviation in parentheses. Across domains and models, matching produces small differences in individual belief probability while retaining substantially larger differences in graded stability.
Dataset
Model
Mean ΔM
Bootstrap SE
95% CI
City Locations
llama-3.2-3b
−0.0256
0.0118
[−0.0488,−0.0018]
llama-3.1-8b
0.0220
0.0308
[−0.0402,0.0815]
llama-3.1-70b
0.0001
0.0049
[−0.0098,0.0095]
gemma-7b
0.0101
0.0080
[−0.0058,0.0256]
gemma-2-9b
0.0180
0.0077
[0.0030,0.0329]
gemma-2-27b
0.0784
0.0144
[0.0497,0.1069]
Appendix
Table A8: Uncertainty in behavioral resilience effects. We report the mean behavioral movement difference ΔM , bootstrap standard error, and 95% percentile bootstrap confidence interval using the sAwMIL probe and Direct Conditional graded stability. Uncertainty is estimated from 10,000 sequence-stratified matched-pair bootstrap resamples. Mean ΔM is positive in 83.3% of model–domain settings; the 95% confidence interval is entirely positive in 30.6% of settings, entirely negative in 2.8% , and includes zero in the remaining 66.7% .
City Locations
Medical Indications
Word Definitions
Model
SVM
Mass Mean
SVM
Mass Mean
SVM
Mass Mean
llama-3.2-3b
11
7
12
11
11
11
llama-3.1-8b
25
17
14
13
13
14
llama-3.1-70b
76
36
65
37
29
30
gemma-7b
17
19
17
17
16
16
gemma-2-9b
22
23
20
21
20
19
Appendix
Table A9 : Selected layers for the alternative probes. Zero-indexed layers selected independently for the SVM and Mass Mean probes for each instruction-tuned LLM. The selected probe-specific layer is subsequently used for the corresponding individual belief probability and Direct Conditional analyses.
Data Science Institute University of Chicago Chicago, USA · Department of Computer Science University of Chicago Chicago, USA · Booth School of Business University of Chicago Chicago, USA