Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer the questions. Existing faithfulness evaluations typically assess each question-context instance in isolation; however, such instance-level evaluation fails to capture a fundamental requirement of faithful behavior: the ability to adapt model responses to changes in available contexts. In particular, a model should provide correct answers when sufficient evidence is present and abstain when it is not. In this work, we propose a Pairwise Faithfulness Benchmark (PFaithBench) that evaluates whether a model can switch between answering and abstaining for the same question under supporting versus non-supporting contexts. Our evaluations across thirty-nine models with seven model families demonstrate that faithfulness fundamentally involves a trade-off between answering and abstaining, and that most current models exhibit a strong bias toward answering, with most faithfulness errors arising from over-answering, i.e., models tend to fabricate a response even when the provided context is insufficient. We further conduct a series of studies on faithfulness training under different data constructions. Our results show that training outcomes are highly sensitive to the specific composition of answering and abstaining data. Constructing answering and abstaining data from mismatched sources can cause models to rely on dataset-specific shortcuts rather than actual context sufficiency. Moreover, increasing answer-supervised data improves answering performance but exacerbates over-answering, while increasing abstaining data reduces hallucination but leads to over-abstention. The code and data are released at https://github.com/tmlr-group/PFaithBench.
Figures & tables
Figure 1: Overview of pairwise faithfulness behavior . Left: Given the same question with sufficient and insufficient contexts, a faithful model should exhibit context-sensitive behavior: answering when sufficient evidence is available and abstaining when key information is missing. Over-abstention and over-answering are two sources of faithfulness errors. Right: Evaluations across seven families with thirty-nine models: most models exhibit bias toward over-answering, where one point is one model.
Figure 2: Pairwise faithfulness evaluation and PFaithBench construction. Left: Existing evaluation protocols assess individual question-context instances, confounding question and context variations. Middle: Pairwise evaluation keeps the question fixed while varying contextual sufficiency, enabling controlled assessment of faithfulness switching. Right: PFaithBench is constructed from multi-hop, single-hop, and counterfactual QA by modifying key information in sufficient contexts.
Family
LLM
PFB-MultiHop
PFB-SingleHop
PFB-ConOri
PFB-ConCF
Avg.
OverAns ↓
OverAbs ↓
OverAns ↓
OverAbs ↓
OverAns ↓
OverAbs ↓
OverAns ↓
OverAbs ↓
OverAns ↓
OverAbs ↓
Llama3
Llama-3.2-1B-Instruct
53.90
35.50
52.76
40.81
44.25
37.04
44.37
43.11
48.82
39.12
Llama-3.2-3B-Instruct
37.40
23.90
12.99
43.18
19.47
23.51
21.11
37.42
22.74
32.00
Llama-3.1-8B-Instruct
56.60
6.50
45.41
6.04
45.89
3.79
47.03
18.46
48.73
8.70
Llama-3.1-70B-Instruct
37.20
7.40
21.00
5.64
28.19
7.71
30.47
18.58
29.21
9.83
Llama-3.3-70B-Instruct
37.20
7.20
27.69
3.54
32.24
6.45
33.12
13.78
32.56
7.74
Table 1: Evaluation of Over-answering and Over-abstention on PFaithBench across diverse LLMs. Lightred indicates OverAns exceeds OverAbs by more than 5%, lightblue indicates OverAbs exceeds OverAns by more than 5%, and lightorange indicates an absolute difference within 5%.
Figure 3: Internal knowledge leakage across models. Leakage measures the extent to which models produce answers by relying on internal parametric knowledge rather than the provided context. Colors indicate different model families, and model sizes increase from left to right within each family.
Family
LLM
PFB-MultiHop
PFB-SingleHop
PFB-ConOri
PFB-ConCF
Avg.
DFAcc ↑
SFAcc ↑
DFAcc ↑
SFAcc ↑
DFAcc ↑
SFAcc ↑
DFAcc ↑
SFAcc ↑
DFAcc ↑
SFAcc ↑
Gap ↓
Llama3
Llama-3.2-1B-Instruct
21.50
5.60
17.06
6.69
27.81
7.21
26.68
3.79
23.26
5.82
17.44
Llama-3.2-3B-Instruct
42.20
26.00
46.19
25.85
59.29
43.87
47.91
15.93
48.90
27.91
20.99
Llama-3.1-8B-Instruct
37.70
30.00
49.87
39.50
50.95
45.13
42.10
19.09
45.15
33.43
11.72
Llama-3.1-70B-Instruct
56.10
50.00
73.62
68.90
64.85
59.04
54.61
22.50
62.30
50.11
12.19
Llama-3.3-70B-Instruct
56.00
49.30
69.16
64.70
62.20
57.77
55.25
26.30
60.65
49.52
11.13
Table 2: Evaluation of pairwise accuracy DFAcc and SFAcc on PFaithBench across diverse LLMs. The Gap=DFAcc−SFAcc indicates the discrepancy between appropriate answer-or-abstain decisions and correct answering. Lightred indicates large gaps (Gap ≥10 ), while lightblue indicates small gaps (Gap <10 ). The best results in each family are highlighted in bold .
Family
LLM
PFB-MultiHop
PFB-SingleHop
PFB-ConOri
PFB-ConCF
Avg.
AnsAcc ↑
AbsAcc ↑
AnsAcc ↑
AbsAcc ↑
AnsAcc ↑
AbsAcc ↑
AnsAcc ↑
AbsAcc ↑
AnsAcc ↑
AbsAcc ↑
Llama3
Llama-3.2-1B-Instruct
18.60
46.10
21.92
47.24
21.49
55.75
6.19
55.63
17.05
51.18
Llama-3.2-3B-Instruct
45.90
62.60
32.28
87.01
57.65
80.53
18.20
78.89
38.51
77.26
Llama-3.1-8B-Instruct
73.80
43.40
73.36
54.59
84.45
54.11
29.58
52.97
65.30
51.27
Llama-3.1-70B-Instruct
82.10
62.80
86.88
79.00
85.46
71.81
30.34
69.53
71.19
70.79
Llama-3.3-70B-Instruct
81.60
62.80
88.71
72.31
86.35
67.76
34.51
66.88
72.79
67.44
Table 3: Evaluation of AnsAcc and AbsAcc on PFaithBench across diverse LLMs. Lightred indicates that the absolute difference ∣AnsAcc−AbsAcc∣ is larger than 10%, while lightblue indicates that the absolute difference is at most 10%. The best results in each family are highlighted in bold .
LLM
Setting
PFB-MultiHop
PFB-SingleHop
DFAcc ↑
SFAcc ↑
AnsAcc ↑
AbsAcc ↑
OverAns ↓
OverAbs ↓
DFAcc ↑
SFAcc ↑
AnsAcc ↑
AbsAcc ↑
OverAns ↓
OverAbs ↓
Llama-3.1-8B-Instruct
Base Model
37.70
30.00
73.80
43.40
56.60
6.50
49.87
39.50
73.36
54.59
45.41
6.04
MHSuff-SHInsuff
34.20
27.40
81.80
36.60
63.40
2.60
37.14
31.50
32.28
98.56
1.44
61.68
SHSuff-MHInsuff
2.70
0.00
0.00
96.70
3.30
95.60
16.01
14.70
91.60
16.80
83.20
1.05
PairContext
71.50
61.00
76.40
81.90
18.10
10.60
79.13
75.33
86.22
86.35
13.65
7.74
Llama-3.2-3B-Instruct
Base Model
42.20
26.00
45.90
62.60
37.40
23.90
46.19
25.85
32.28
87.01
12.99
43.18
Table 4: Impact of context coupling on faithfulness training. The best value of each metric is bold . Among the three training settings, lightblue marks both AnsAcc and AbsAcc for the smallest ∣AnsAcc−AbsAcc∣ , and marks both OverAns and OverAbs for the smallest ∣OverAns−OverAbs∣ .
Figure 4: Impact of answering–abstention data composition (asymmetric reward). Faithfulness training under different answering data ratios α∈{0,25,50,75,100} across three LLMs. This is the average value across four subsets in PFaithBench. More results are provided in Appendix C.2 .
Setting
AMC
AIME24
AIME25
MATH500
MMLU-Pro
GPQA-D
IFEval
Setting
AMC
AIME24
AIME25
MATH500
MMLU-Pro
GPQA-D
IFEval
Llama-3.1-8B-Instruct
α=0
16.42
3.02
0.42
31.80
34.21
26.26
78.82
Base Model
14.76
2.19
0.63
33.40
36.10
29.29
76.86
α=25
17.32
3.33
0.21
34.60
36.43
30.81
77.42
MHSuff-SHInsuff
19.28
3.96
0.52
38.60
35.61
25.76
80.34
α=50
18.37
3.33
0.52
34.60
36.18
25.25
78.87
SHSuff-MHInsuff
18.52
2.71
0.31
36.20
35.46
35.35
79.53
α=75
17.62
3.23
0.31
37.60
35.85
27.78
78.53
PairContext
19.43
4.17
0.31
38.00
36.58
30.30
77.81
α=100
18.52
4.17
0.31
36.80
36.58
27.78
78.39
Llama-3.2-3B-Instruct
α=0
15.06
5.42
0.10
34.60
28.99
26.77
74.14
Table 5: Impact of faithfulness training on general model capabilities.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Evaluation Paradigm
Answerable
Unanswerable
Counterfactual
Reasoning Type
FaithEval ( Ming et al., 2025 )
Instance-level
✓
✓
✓
Single-hop
ConFiQA ( Bi et al., 2025 )
Instance-level
✓
✓
Single-hop
TimeQA ( Chen et al., 2021 )
Instance-level
✓
✓
Multi-hop
RGB ( Chen et al., 2024 )
Instance-level
✓
Single-hop
ClashEval ( Wu et al., 2024 )
Instance-level
✓
Multi-hop
SQuAD ( Rajpurkar et al., 2016 )
Instance-level
✓
Single-hop
Appendix
Table 6: Relationship between PFaithBench and existing benchmarks.
Subset
# Eval Samples
# Training Samples
Data Sources
PFB-MultiHop
1,000
2,000
HotpotQA & 2WikiMultiHopQA
PFB-SingleHop
762
1,997
SQuAD
PFB-ConOri
791
–
ConFiQA
PFB-ConCF
791
–
ConFiQA
Total
3,344
3,997
HotpotQA, 2WikiQA, SQuAD, ConFiQA
Appendix
Table 7: Statistics of PFaithBench.
Figure 5: Example of sufficient and insufficient contexts in PFaithBench.
Family
LLM
PFB-MultiHop
PFB-SingleHop
PFB-ConOri
PFB-ConCF
Avg.
Imbalance ↓
Asymmetry ↓
Imbalance ↓
Asymmetry ↓
Imbalance ↓
Asymmetry ↓
Imbalance ↓
Asymmetry ↓
Imbalance ↓
Asymmetry ↓
Llama3
Llama-3.2-1B-Instruct
27.50
40.50
25.33
40.55
34.26
48.55
49.43
51.83
34.13
45.36
Llama-3.2-3B-Instruct
16.70
36.60
54.72
61.15
22.88
36.66
60.68
62.96
38.75
49.34
Llama-3.1-8B-Instruct
30.40
43.80
18.77
33.86
30.34
39.32
23.39
33.88
25.72
37.71
Llama-3.1-70B-Instruct
19.30
32.10
7.87
17.98
13.65
26.42
39.19
47.03
20.00
30.88
Llama-3.3-70B-Instruct
18.80
32.30
16.40
24.02
18.58
28.57
32.36
40.58
21.54
31.37
Appendix
Table 8: Imbalance and asymmetry on PFaithBench across diverse LLMs. Imbalance is ∣AnsAcc−AbsAcc∣ ; asymmetry is ∣max(AnsAcc,AbsAcc)−SFAcc∣ . Avg. is the mean across four subsets. The minimum in each model family and column is highlighted in bold .
Figure 6: Impact of answering–abstention data composition (symmetric reward). Faithfulness training under different answering data ratios α∈{0,25,50,75,100} across three LLMs. This is the average value across four subsets in PFaithBench.
Model Family
Avg. Model Size (B)
SFAcc ↑
OverAns ↓
OverAbs ↓
AnsAcc ↑
AbsAcc ↑
Leakage ↓
Llama
30.40 Billion
33.36
36.41
19.48
52.97
63.59
14.67
Qwen3 (NoThink)
10.05 Billion
42.64
33.73
14.52
64.82
66.27
10.34
Qwen3 (Think)
10.05 Billion
31.09
59.66
3.67
74.90
40.34
14.66
Qwen3.5 (NoThink)
8.56 Billion
45.47
35.28
7.06
72.05
64.72
8.80
Qwen3.5 (Think)
8.56 Billion
39.58
46.57
5.06
74.98
53.43
6.92
Qwen3.6
27.00 Billion
44.47
42.21
4.26
79.98
57.79
12.67
Appendix
Table 9: Family-level performance across base models on PFaithBench. Metrics are averaged equally across models and the four subsets. Average model sizes are in billions of parameters.
LLM
α
PFB-MultiHop
PFB-SingleHop
DFAcc ↑
SFAcc ↑
OverAns ↓
OverAbs ↓
Leakage ↓
DFAcc ↑
SFAcc ↑
OverAns ↓
OverAbs ↓
Leakage ↓
Llama-3.1-8B-Instruct
α=0
10.90
0.00
10.50
84.70
0.00
18.50
0.00
2.23
80.05
0.00
α=25
60.60
47.90
28.80
12.10
9.10
76.25
65.09
9.58
14.96
0.79
α=50
69.10
59.10
17.80
13.60
10.40
76.51
73.36
13.25
10.50
1.44
α=75
58.50
49.40
36.10
5.90
18.50
62.99
59.71
34.51
3.02
2.89
α=100
0.40
0.40
99.60
0.00
52.80
0.26
0.26
99.74
0.00
12.73
Appendix
Table 10: Detailed impact of answering-abstention data composition (asymmetric reward). The answering data percentage α vary α∈{0,25,50,75,100} , serving as a supplement to Fig. 4 .
LLM
α
PFB-MultiHop
PFB-SingleHop
DFAcc ↑
SFAcc ↑
OverAns ↓
OverAbs ↓
Leakage ↓
DFAcc ↑
SFAcc ↑
OverAns ↓
OverAbs ↓
Leakage ↓
Llama-3.1-8B-Instruct
α=0
0.00
0.00
0.00
100.00
0.00
0.00
0.00
0.00
100.00
0.00
α=25
17.00
0.30
9.30
78.90
0.00
24.67
0.52
1.97
74.67
0.00
α=50
68.10
58.70
17.10
15.40
11.90
78.35
72.97
11.68
10.24
1.44
α=75
63.90
53.60
30.00
6.30
18.90
71.39
67.59
26.38
2.76
3.02
α=100
2.70
2.10
97.30
0.10
51.50
3.67
3.54
96.33
0.00
9.32
Appendix
Table 11: Detailed impact of answering-abstention data composition (symmetric reward). The answering data percentage α vary α∈{0,25,50,75,100} , serving as a supplement to Fig. 6 .
Setting
Avg.
Prompt Strict
Inst. Strict
Prompt Loose
Inst. Loose
Setting
Prompt Strict
Inst. Strict
Prompt Loose
Inst. Loose
Llama-3.1-8B-Instruct
α=0
78.82
73.20
81.06
77.08
83.93
Base Model
76.86
70.24
79.38
74.86
82.97
α=25
77.42
71.90
79.74
75.42
82.61
MHSuff-SHInsuff
80.34
74.49
82.49
78.74
85.61
α=50
78.87
72.64
81.06
77.26
84.53
SHSuff-MHInsuff
79.53
74.31
82.01
77.26
84.53
α=75
78.53
72.46
80.70
76.89
84.05
PairContext
77.81
71.35
80.22
75.97
83.69
α=100
78.39
72.27
80.82
76.52
83.93
Llama-3.2-3B-Instruct
α=0
74.14
67.65
76.38
72.09
80.46
Appendix
Table 12: Detailed results on IFEval. IFEval reports four metrics: prompt-level strict accuracy, prompt-level loose accuracy, instruction-level strict accuracy, and instruction-level loose accuracy. In Table 5 , we report the average of these four metrics. Here, we provide the detailed results.
Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an underutilized behavioral property of LLMs: a credible answer should remain stable when the same task is placed under topic-aligned, content-neutral contextual variation. Building on this intuition, we operationalize Cross-Contextual Consistency (C3) by comparing model generations under original and perturbed prompts. Across 26 models and six benchmarks spanning reasoning, factuality, and code generation, we find that answers with smaller cross-contextual shifts are more likely to be correct or factual. We demonstrate that C3 provides a complementary axis of evaluation and can serve as a benchmark usefulness diagnostic, identifying which portions of a benchmark remain informative even when aggregated scores are widely considered "saturate".
Siyang Wu, Yibo Jiang, Bryon Aragam
Data Science Institute University of Chicago Chicago, USA · Department of Computer Science University of Chicago Chicago, USA · Booth School of Business University of Chicago Chicago, USA
Large language models (LLMs) are prone to hallucinating or misinterpreting facts, which impairs their usability in retrieval-augmented generation or data-to-text systems. We analyse how faithfulness of LLMs to provided context depends on how plausible they perceive the context to be (context-memory conflict). To better identify error patterns, we make use of the increased difficulty of non-English and low-resource language text generation and input data based on local knowledge, only partially captured in models' parametric knowledge. We let the models generate text in English, Czech, Slovak and Upper Sorbian from factual (FA), counterfactual (CFA) and fictional (FI) RDF triples containing local Czech and Slovak data. Contrary to our expectations, we observe only a weak context-memory conflict on the human-annotated sample. For Kimi K3 as an LLM judge, which agrees well with human annotations on the sample, counterfactual inputs receive only slightly lower faithfulness scores than factual ones (-0.05 on a 1-5 scale). We also find that a suboptimal choice of LLM judge would lead to overestimating the strength of the context-memory conflict.
Peter Kochelka, Aleš Manuel Papáček, Vojtěch Dvořák +1
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics, Prague, Czech Republic
Chains of thought (CoTs) have become central in interpreting and auditing behaviors of large language models. Yet growing evidence suggests that these traces often fail to faithfully represent the computations behind a model's predictions. Several faithfulness metrics have been proposed, but whether they indeed measure faithfulness remains unknown. Answering this requires ground-truth labels, which are hard to obtain since internal computations are not directly observable. Consequently, most works proposing metrics report only absolute scores or comparisons to prior metrics, and the few existing benchmarks rely on proxies like plausibility or importance, properties orthogonal to faithfulness that can mislead about whether a CoT can be trusted. We address this challenge by constructing tasks whose outputs reveal which intermediate computations must have produced them, and developing an automated labeling pipeline that yields ground-truth faithfulness labels at both the step and CoT level. Building on this methodology, we present BonaFide, a benchmark of 3,066 labeled CoTs across 13 tasks and 10 models, and use it to conduct the first systematic evaluation of prominent faithfulness metrics. Our experiments show that most metrics perform near chance, exhibit strong prediction biases and degrade on longer CoTs. The best metric reaches only 0.70 AUROC at the CoT level while another reaches 0.59 at the step level, with neither transferring across settings, while entailing prohibitively high computational cost. Our results expose fundamental gaps in current faithfulness evaluation and call for the development of more reliable and efficient metrics.