Large language models (LLMs) increasingly mediate how people access and reason with information, yet factual reliability is usually evaluated one judgment at a time. We introduce graded belief stability, a relational measure of how well a belief persists within an LLM's broader belief system. Unlike individual belief probability, it asks whether support for a claim persists when that claim is considered alongside the model's other epistemic commitments. We operationalize this idea with a Direct Conditional estimator that uses internal model representations to estimate conditional belief probabilities. Across 12 LLMs and three domains, lower-stability beliefs exhibit greater mean behavioral movement under conversational challenge in 83.3% of model-domain settings after matching on individual belief probability. Graded belief stability therefore extends reliability assessment beyond how strongly an LLM supports a claim to how robustly that belief is supported within its broader system of beliefs.
Figures & tables
Official Name
Short Name
# Layers
# Parameters
Release Date
Source
Citation
Gemma- 7 b-it
gemma-7b
28
8.54 B
Feb 21 , 2024
Google
[ 22 ]
Gemma- 2 - 9 b-it
gemma-2-9b
42
9.24 B
Jun 27 , 2024
Google
[ 22 ]
Gemma- 2 - 27 b-it
gemma-2-27b
46
27.23 B
Jun 27 , 2024
Google
[ 22 ]
Llama- 3.2 - 3 b-Instruct
llama-3.2-3b
28
3.21 B
Sep 25 , 2024
Meta
[ 23 ]
Llama- 3.1 - 8 b-Instruct
llama-3.1-8b
32
8.03 B
Jul 23 , 2024
Meta
[ 23 ]
Llama- 3.1 - 70 b-Instruct
llama-3.1-70b
80
70.55 B
Jul 23 , 2024
Meta
[ 23 ]
Table 1 : LLMs used in graded stability experiments. We list the official names of the LLMs according to the HuggingFace repository [ 45 ] , where they are publicly available. We further specify the shortened name used throughout the paper, the number of layers, parameter count, release date, source organization, and official citation.
Dataset
True
False
Synthetic
Examples
City Locations
A: 1392 N: 1376
A: 1358 N: 1374
A: 876 N: 876
T. The city of Surat is located in India. F. The city of Palembang is located in the Dominican Republic. S. The city of Norminsk is located in Jamoates.
Medical Indications
A: 1439 N: 1522
A: 1523 N: 1419
A: 478 N: 522
T. Pentobarbital is indicated for the treatment of insomnia. F. Vancomycin is not indicated for the treatment of lower respiratory tract infections. S. Alumil is indicated for the treatment of reticers.
Word Definitions
A: 1234 N: 1235
A: 1277 N: 1254
A: 1747 N: 1753
T. Hoagy is a synonym of an Italian sandwich. F. Decalogue is an astronomer. S. Dostab is a scencer.
Table 2 : Summary of datasets and statement types. Number of affirmative (A) and negated (N) statements across the three domains, along with examples. Each dataset includes True (T), False (F), and Synthetic (S) statements. Synthetic statements serve as Neither statements constructed to minimize prior LLM exposure. An earlier version of this table was introduced in [ 38 ] .
si
sj
si→sj
T
T
T
T
N
N
T
F
F
N
T
T
N
N
N
N
F
F
Table 3 : Cooper–Cantwell conditional truth table. We report whether the conditional si→sj≡sj∣si is True (T), False (F), or Neither (N) based on the truth values of si and sj according to the Cooper–Cantwell conditional [ 11 , 8 ] . When the antecedent si is True or Neither , the conditional takes the truth value of the consequent sj .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Description
M
A fixed large language model (LLM).
s
Natural-language statement.
P
A proposition believed by M whose stability is evaluated.
x
A non-disbelieved proposition used as conditioning information.
c∈{T,F,N}
Trivalent veracity class: True ( T ), False ( F ), or Neither ( N ).
K
Number of probe classes.
Appendix
Table A1 : Notation. Summary of the principal mathematical symbols used throughout the manuscript.
Dataset
Train
Calibration
Test
Total
City Locations
3999(0.55)
1398(0.19)
1855(0.26)
7252(1.00)
Medical Indications
3849(0.56)
1327(0.19)
1727(0.25)
6903(1.00)
Word Definitions
4717(0.55)
1628(0.19)
2155(0.25)
8500(1.00)
Appendix
Table A2 : Dataset splits. Number of statements used for training, calibration, and testing. Proportions of the full dataset are reported in parentheses. A version of this table appears in [ 38 ] .
Official Name
Short Name
# Layers
# Parameters
Release Date
Source
Citation
Gemma- 7 b
gemma-7b (b)
28
8.54 B
Feb 21 , 2024
Google
[ 22 ]
Gemma- 2 - 9 b
gemma-2-9b (b)
42
9.24 B
Jun 27 , 2024
Google
[ 22 ]
Gemma- 2 - 27 b
gemma-2-27b (b)
46
27.23 B
Jun 27 , 2024
Google
[ 22 ]
Llama- 3.2 - 3 b
llama-3.2-3b (b)
28
3.21 B
Sep 25 , 2024
Meta
[ 23 ]
Llama- 3.1 - 8 b
llama-3.1-8b (b)
32
8.03 B
Jul 23 , 2024
Meta
[ 23 ]
Llama- 3.1 - 70 b
llama-3.1-70b (b)
80
70.55 B
Jul 23 , 2024
Meta
[ 23 ]
Appendix
Table A3 : Base LLMs used in graded stability experiments. We list the official names of the LLMs according to the HuggingFace repository [ 45 ] , where they are publicly available. We further specify the shortened name used throughout the paper, the number of layers, parameter count, release date, source organization, and official citation.
Model
City Locations
Medical Indications
Word Definitions
llama-3.2-3b
9
11
13
llama-3.1-8b
12
19
12
llama-3.1-70b
34
34
17
gemma-7b
19
16
17
gemma-2-9b
23
20
21
gemma-2-27b
15
19
18
Appendix
Table A4 : Selected sAwMIL layers. Zero-indexed layers selected independently for each instruction-tuned LLM and dataset by minimizing three-class calibration log loss. The selected layer is subsequently used for the corresponding individual-statement, Direct Conditional, and Joint sAwMIL probes.
Belief sets
Conditional-pair coverage
Dataset
Model
∣BM∣
∣XM∣
tM
Candidate
Direct
Joint
Joint excl.
City Locations
llama-3.2-3b
632
1,217
0.506
768,512
768,512
768,512
0(0.0000%)
llama-3.1-8b
629
1,213
0.543
762,348
762,348
762,348
0(0.0000%)
llama-3.1-70b
628
1,212
0.478
760,508
760,508
760,471
37(0.0049%)
gemma-7b
615
1,199
0.513
736,770
736,770
736,770
0(0.0000%)
gemma-2-9b
626
1,211
0.537
757,460
757,460
756,727
733(0.0968%)
Appendix
Table A5: Belief-set and conditional-pair coverage using sAwMIL . For each LLM and domain, ∣BM∣ and ∣XM∣ denote the numbers of beliefs and non-disbeliefs, respectively, and tM is the empirical belief threshold. Candidate is the number ∣BM∣(∣XM∣−1) of constructed (P,x) pairs after excluding self-pairs. Direct and Joint report the numbers of valid conditional-probability estimates. Joint excl. reports the number and percentage of candidate pairs excluded because the Joint-to-Conditional denominator is at most 10−12 .
Dataset
Model
Exact
πTdirect(P,x)>πT(P)
πFdirect(P,x)>πF(P)
πF(x)>πNdirect(P,x)
πNdirect(P,x)>πN(P)+πF(x)
City Locations
llama-3.2-3b
0.000%
2.7%(0.095)
90.4%(0.154)
3.4%(0.122)
96.6%(0.207)
llama-3.1-8b
0.000%
5.2%(0.054)
88.0%(0.051)
5.8%(0.047)
94.2%(0.080)
llama-3.1-70b
<0.001%
7.9%(0.031)
80.5%(0.035)
11.4%(0.026)
88.5%(0.047)
gemma-7b
<0.001%
3.8%(0.107)
92.4%(0.159)
7.9%(0.091)
92.1%(0.102)
gemma-2-9b
0.000%
3.1%(0.050)
93.9%(0.087)
11.0%(0.036)
89.0%(0.046)
gemma-2-27b
0.000%
2.4%(0.067)
93.7%(0.104)
5.9%(0.073)
94.1%(0.109)
Appendix
Table A6: Exact CCK coherence and constraint violations for Direct Conditional estimates using sAwMIL . For each LLM and domain, we report the percentage of (P,x) pairs satisfying all four CCK constraints simultaneously within tolerance 10−8 . Here, 0.000% signifies exactly zero fully coherent pairs, while <0.001% means a small number of exactly coherent pairs exists. Each remaining column reports the percentage of pairs violating each constraint, with the mean excess across violating pairs shown in parentheses. Exact coherence is rare across all models and domains, with no model–domain combination exceeding 0.206% . Violations occur most frequently for the constraints on conditional False and Neither probability mass.
Dataset
Model
n
∣ΔπT∣
∣Δγ∣
City Locations
llama-3.2-3b
315
0.00435(0.01055)
0.235(0.173)
llama-3.1-8b
305
0.00181(0.00556)
0.090(0.122)
llama-3.1-70b
287
0.00483(0.02893)
0.066(0.116)
gemma-7b
304
0.00377(0.01042)
0.175(0.175)
gemma-2-9b
291
0.00370(0.01686)
0.082(0.141)
gemma-2-27b
316
0.00362(0.01372)
0.133(0.143)
Appendix
Table A7: Behavioral matching quality. For each LLM and domain, n is the number of matched proposition pairs, ∣ΔπT∣ is the absolute difference in belief probability between the two propositions in each pair, and ∣Δγ∣ is their absolute difference in graded stability. Values for ∣ΔπT∣ and ∣Δγ∣ report the mean with standard deviation in parentheses. Across domains and models, matching produces small differences in individual belief probability while retaining substantially larger differences in graded stability.
Dataset
Model
Mean ΔM
Bootstrap SE
95% CI
City Locations
llama-3.2-3b
−0.0256
0.0118
[−0.0488,−0.0018]
llama-3.1-8b
0.0220
0.0308
[−0.0402,0.0815]
llama-3.1-70b
0.0001
0.0049
[−0.0098,0.0095]
gemma-7b
0.0101
0.0080
[−0.0058,0.0256]
gemma-2-9b
0.0180
0.0077
[0.0030,0.0329]
gemma-2-27b
0.0784
0.0144
[0.0497,0.1069]
Appendix
Table A8: Uncertainty in behavioral resilience effects. We report the mean behavioral movement difference ΔM , bootstrap standard error, and 95% percentile bootstrap confidence interval using the sAwMIL probe and Direct Conditional graded stability. Uncertainty is estimated from 10,000 sequence-stratified matched-pair bootstrap resamples. Mean ΔM is positive in 83.3% of model–domain settings; the 95% confidence interval is entirely positive in 30.6% of settings, entirely negative in 2.8% , and includes zero in the remaining 66.7% .
City Locations
Medical Indications
Word Definitions
Model
SVM
Mass Mean
SVM
Mass Mean
SVM
Mass Mean
llama-3.2-3b
11
7
12
11
11
11
llama-3.1-8b
25
17
14
13
13
14
llama-3.1-70b
76
36
65
37
29
30
gemma-7b
17
19
17
17
16
16
gemma-2-9b
22
23
20
21
20
19
Appendix
Table A9 : Selected layers for the alternative probes. Zero-indexed layers selected independently for the SVM and Mass Mean probes for each instruction-tuned LLM. The selected probe-specific layer is subsequently used for the corresponding individual belief probability and Direct Conditional analyses.
Large Language Models (LLMs) are increasingly employed in various question-answering tasks. However, recent studies showcase that LLMs are susceptible to persuasion and could adopt counterfactual beliefs. We present a systematic evaluation of LLM susceptibility to persuasion under the \emph{Source--Message--Channel--Receiver} (SMCR) communication framework. Across six mainstream Large Language Models (LLMs) and three domains (factual knowledge, medical QA, and social bias), we analyze how different persuasive strategies influence stated belief stability over multiple interaction turns. We further examine whether verbalized confidence prompting (i.e., eliciting self-reported confidence scores) affects resistance to persuasion. Results show that the smallest model (Llama 3.2-3B) exhibits extreme compliance, with 82.5% of belief changes occurring at the first persuasive turn (average end turn of 1.1--1.4). Contrary to expectations, verbalized confidence prompting \emph{increases} vulnerability by accelerating belief erosion rather than enhancing robustness. Finally, an exploratory study of adversarial fine-tuning reveals highly model-dependent effectiveness: GPT-4o-mini achieves near-complete robustness (98.6%), and Mistral~7B improves substantially (35.7% → 79.3%), but Llama models remain highly susceptible (<14% RQ1) even when fine-tuned on their own failure cases. Together, these findings highlight substantial model-dependent limits of current robustness interventions and offer guidance for developing more trustworthy LLMs.
Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLMs), as they are increasingly deployed in user-facing settings. Prior work showed that even capable LLMs exhibit a systemic weakness in acknowledging user beliefs grounded in incorrect information. We extend this evaluation to 10 LLMs across 18 epistemic expressions and find that the size and direction of this weakness depend on the verb used to express the belief, with the accuracy gap between factual and false information ranging from +50% on "I vaguely remember" to -14% on "I seriously doubt". We further show that the phenomenon stems from what we call task confusion: models default to fact-checking the underlying claim, overriding the user's stated belief. We provide evidence where chains of thought that explicitly fact-check show lower accuracy on false information than those that do not, and a single instruction can reverse the failure across verb families. Mechanistically, models attend more to false beliefs they fail to confirm, but suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods. Our findings clarify prior results and show how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs.
Quang Minh Nguyen, Luis Frentzen Salim
1KAIST · 2National Taiwan University of Science and Technology
Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an underutilized behavioral property of LLMs: a credible answer should remain stable when the same task is placed under topic-aligned, content-neutral contextual variation. Building on this intuition, we operationalize Cross-Contextual Consistency (C3) by comparing model generations under original and perturbed prompts. Across 26 models and six benchmarks spanning reasoning, factuality, and code generation, we find that answers with smaller cross-contextual shifts are more likely to be correct or factual. We demonstrate that C3 provides a complementary axis of evaluation and can serve as a benchmark usefulness diagnostic, identifying which portions of a benchmark remain informative even when aggregated scores are widely considered "saturate".
Siyang Wu, Yibo Jiang, Bryon Aragam
Data Science Institute University of Chicago Chicago, USA · Department of Computer Science University of Chicago Chicago, USA · Booth School of Business University of Chicago Chicago, USA