Large language models (LLMs) are expected to answer questions faithfully based on the provided context, abstaining when the context information is insufficient to answer the questions. Existing faithfulness evaluations typically assess each question-context instance in isolation; however, such instance-level evaluation fails to capture a fundamental requirement of faithful behavior: the ability to adapt model responses to changes in available contexts. In particular, a model should provide correct answers when sufficient evidence is present and abstain when it is not. In this work, we propose a Pairwise Faithfulness Benchmark (PFaithBench) that evaluates whether a model can switch between answering and abstaining for the same question under supporting versus non-supporting contexts. Our evaluations across thirty-nine models with seven model families demonstrate that faithfulness fundamentally involves a trade-off between answering and abstaining, and that most current models exhibit a strong bias toward answering, with most faithfulness errors arising from over-answering, i.e., models tend to fabricate a response even when the provided context is insufficient. We further conduct a series of studies on faithfulness training under different data constructions. Our results show that training outcomes are highly sensitive to the specific composition of answering and abstaining data. Constructing answering and abstaining data from mismatched sources can cause models to rely on dataset-specific shortcuts rather than actual context sufficiency. Moreover, increasing answer-supervised data improves answering performance but exacerbates over-answering, while increasing abstaining data reduces hallucination but leads to over-abstention. The code and data are released at https://github.com/tmlr-group/PFaithBench.
Figures & tables
Figure 1: Overview of pairwise faithfulness behavior . Left: Given the same question with sufficient and insufficient contexts, a faithful model should exhibit context-sensitive behavior: answering when sufficient evidence is available and abstaining when key information is missing. Over-abstention and over-answering are two sources of faithfulness errors. Right: Evaluations across seven families with thirty-nine models: most models exhibit bias toward over-answering, where one point is one model.
Figure 2: Pairwise faithfulness evaluation and PFaithBench construction. Left: Existing evaluation protocols assess individual question-context instances, confounding question and context variations. Middle: Pairwise evaluation keeps the question fixed while varying contextual sufficiency, enabling controlled assessment of faithfulness switching. Right: PFaithBench is constructed from multi-hop, single-hop, and counterfactual QA by modifying key information in sufficient contexts.
Family
LLM
PFB-MultiHop
PFB-SingleHop
PFB-ConOri
PFB-ConCF
Avg.
OverAns ↓
OverAbs ↓
OverAns ↓
OverAbs ↓
OverAns ↓
OverAbs ↓
OverAns ↓
OverAbs ↓
OverAns ↓
OverAbs ↓
Llama3
Llama-3.2-1B-Instruct
53.90
35.50
52.76
40.81
44.25
37.04
44.37
43.11
48.82
39.12
Llama-3.2-3B-Instruct
37.40
23.90
12.99
43.18
19.47
23.51
21.11
37.42
22.74
32.00
Llama-3.1-8B-Instruct
56.60
6.50
45.41
6.04
45.89
3.79
47.03
18.46
48.73
8.70
Llama-3.1-70B-Instruct
37.20
7.40
21.00
5.64
28.19
7.71
30.47
18.58
29.21
9.83
Llama-3.3-70B-Instruct
37.20
7.20
27.69
3.54
32.24
6.45
33.12
13.78
32.56
7.74
Table 1: Evaluation of Over-answering and Over-abstention on PFaithBench across diverse LLMs. Lightred indicates OverAns exceeds OverAbs by more than 5%, lightblue indicates OverAbs exceeds OverAns by more than 5%, and lightorange indicates an absolute difference within 5%.
Figure 3: Internal knowledge leakage across models. Leakage measures the extent to which models produce answers by relying on internal parametric knowledge rather than the provided context. Colors indicate different model families, and model sizes increase from left to right within each family.
Family
LLM
PFB-MultiHop
PFB-SingleHop
PFB-ConOri
PFB-ConCF
Avg.
DFAcc ↑
SFAcc ↑
DFAcc ↑
SFAcc ↑
DFAcc ↑
SFAcc ↑
DFAcc ↑
SFAcc ↑
DFAcc ↑
SFAcc ↑
Gap ↓
Llama3
Llama-3.2-1B-Instruct
21.50
5.60
17.06
6.69
27.81
7.21
26.68
3.79
23.26
5.82
17.44
Llama-3.2-3B-Instruct
42.20
26.00
46.19
25.85
59.29
43.87
47.91
15.93
48.90
27.91
20.99
Llama-3.1-8B-Instruct
37.70
30.00
49.87
39.50
50.95
45.13
42.10
19.09
45.15
33.43
11.72
Llama-3.1-70B-Instruct
56.10
50.00
73.62
68.90
64.85
59.04
54.61
22.50
62.30
50.11
12.19
Llama-3.3-70B-Instruct
56.00
49.30
69.16
64.70
62.20
57.77
55.25
26.30
60.65
49.52
11.13
Table 2: Evaluation of pairwise accuracy DFAcc and SFAcc on PFaithBench across diverse LLMs. The Gap=DFAcc−SFAcc indicates the discrepancy between appropriate answer-or-abstain decisions and correct answering. Lightred indicates large gaps (Gap ≥10 ), while lightblue indicates small gaps (Gap <10 ). The best results in each family are highlighted in bold .
Family
LLM
PFB-MultiHop
PFB-SingleHop
PFB-ConOri
PFB-ConCF
Avg.
AnsAcc ↑
AbsAcc ↑
AnsAcc ↑
AbsAcc ↑
AnsAcc ↑
AbsAcc ↑
AnsAcc ↑
AbsAcc ↑
AnsAcc ↑
AbsAcc ↑
Llama3
Llama-3.2-1B-Instruct
18.60
46.10
21.92
47.24
21.49
55.75
6.19
55.63
17.05
51.18
Llama-3.2-3B-Instruct
45.90
62.60
32.28
87.01
57.65
80.53
18.20
78.89
38.51
77.26
Llama-3.1-8B-Instruct
73.80
43.40
73.36
54.59
84.45
54.11
29.58
52.97
65.30
51.27
Llama-3.1-70B-Instruct
82.10
62.80
86.88
79.00
85.46
71.81
30.34
69.53
71.19
70.79
Llama-3.3-70B-Instruct
81.60
62.80
88.71
72.31
86.35
67.76
34.51
66.88
72.79
67.44
Table 3: Evaluation of AnsAcc and AbsAcc on PFaithBench across diverse LLMs. Lightred indicates that the absolute difference ∣AnsAcc−AbsAcc∣ is larger than 10%, while lightblue indicates that the absolute difference is at most 10%. The best results in each family are highlighted in bold .
LLM
Setting
PFB-MultiHop
PFB-SingleHop
DFAcc ↑
SFAcc ↑
AnsAcc ↑
AbsAcc ↑
OverAns ↓
OverAbs ↓
DFAcc ↑
SFAcc ↑
AnsAcc ↑
AbsAcc ↑
OverAns ↓
OverAbs ↓
Llama-3.1-8B-Instruct
Base Model
37.70
30.00
73.80
43.40
56.60
6.50
49.87
39.50
73.36
54.59
45.41
6.04
MHSuff-SHInsuff
34.20
27.40
81.80
36.60
63.40
2.60
37.14
31.50
32.28
98.56
1.44
61.68
SHSuff-MHInsuff
2.70
0.00
0.00
96.70
3.30
95.60
16.01
14.70
91.60
16.80
83.20
1.05
PairContext
71.50
61.00
76.40
81.90
18.10
10.60
79.13
75.33
86.22
86.35
13.65
7.74
Llama-3.2-3B-Instruct
Base Model
42.20
26.00
45.90
62.60
37.40
23.90
46.19
25.85
32.28
87.01
12.99
43.18
Table 4: Impact of context coupling on faithfulness training. The best value of each metric is bold . Among the three training settings, lightblue marks both AnsAcc and AbsAcc for the smallest ∣AnsAcc−AbsAcc∣ , and marks both OverAns and OverAbs for the smallest ∣OverAns−OverAbs∣ .
Figure 4: Impact of answering–abstention data composition (asymmetric reward). Faithfulness training under different answering data ratios α∈{0,25,50,75,100} across three LLMs. This is the average value across four subsets in PFaithBench. More results are provided in Appendix C.2 .
Setting
AMC
AIME24
AIME25
MATH500
MMLU-Pro
GPQA-D
IFEval
Setting
AMC
AIME24
AIME25
MATH500
MMLU-Pro
GPQA-D
IFEval
Llama-3.1-8B-Instruct
α=0
16.42
3.02
0.42
31.80
34.21
26.26
78.82
Base Model
14.76
2.19
0.63
33.40
36.10
29.29
76.86
α=25
17.32
3.33
0.21
34.60
36.43
30.81
77.42
MHSuff-SHInsuff
19.28
3.96
0.52
38.60
35.61
25.76
80.34
α=50
18.37
3.33
0.52
34.60
36.18
25.25
78.87
SHSuff-MHInsuff
18.52
2.71
0.31
36.20
35.46
35.35
79.53
α=75
17.62
3.23
0.31
37.60
35.85
27.78
78.53
PairContext
19.43
4.17
0.31
38.00
36.58
30.30
77.81
α=100
18.52
4.17
0.31
36.80
36.58
27.78
78.39
Llama-3.2-3B-Instruct
α=0
15.06
5.42
0.10
34.60
28.99
26.77
74.14
Table 5: Impact of faithfulness training on general model capabilities.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Evaluation Paradigm
Answerable
Unanswerable
Counterfactual
Reasoning Type
FaithEval ( Ming et al., 2025 )
Instance-level
✓
✓
✓
Single-hop
ConFiQA ( Bi et al., 2025 )
Instance-level
✓
✓
Single-hop
TimeQA ( Chen et al., 2021 )
Instance-level
✓
✓
Multi-hop
RGB ( Chen et al., 2024 )
Instance-level
✓
Single-hop
ClashEval ( Wu et al., 2024 )
Instance-level
✓
Multi-hop
SQuAD ( Rajpurkar et al., 2016 )
Instance-level
✓
Single-hop
Appendix
Table 6: Relationship between PFaithBench and existing benchmarks.
Subset
# Eval Samples
# Training Samples
Data Sources
PFB-MultiHop
1,000
2,000
HotpotQA & 2WikiMultiHopQA
PFB-SingleHop
762
1,997
SQuAD
PFB-ConOri
791
–
ConFiQA
PFB-ConCF
791
–
ConFiQA
Total
3,344
3,997
HotpotQA, 2WikiQA, SQuAD, ConFiQA
Appendix
Table 7: Statistics of PFaithBench.
Figure 5: Example of sufficient and insufficient contexts in PFaithBench.
Family
LLM
PFB-MultiHop
PFB-SingleHop
PFB-ConOri
PFB-ConCF
Avg.
Imbalance ↓
Asymmetry ↓
Imbalance ↓
Asymmetry ↓
Imbalance ↓
Asymmetry ↓
Imbalance ↓
Asymmetry ↓
Imbalance ↓
Asymmetry ↓
Llama3
Llama-3.2-1B-Instruct
27.50
40.50
25.33
40.55
34.26
48.55
49.43
51.83
34.13
45.36
Llama-3.2-3B-Instruct
16.70
36.60
54.72
61.15
22.88
36.66
60.68
62.96
38.75
49.34
Llama-3.1-8B-Instruct
30.40
43.80
18.77
33.86
30.34
39.32
23.39
33.88
25.72
37.71
Llama-3.1-70B-Instruct
19.30
32.10
7.87
17.98
13.65
26.42
39.19
47.03
20.00
30.88
Llama-3.3-70B-Instruct
18.80
32.30
16.40
24.02
18.58
28.57
32.36
40.58
21.54
31.37
Appendix
Table 8: Imbalance and asymmetry on PFaithBench across diverse LLMs. Imbalance is ∣AnsAcc−AbsAcc∣ ; asymmetry is ∣max(AnsAcc,AbsAcc)−SFAcc∣ . Avg. is the mean across four subsets. The minimum in each model family and column is highlighted in bold .
Figure 6: Impact of answering–abstention data composition (symmetric reward). Faithfulness training under different answering data ratios α∈{0,25,50,75,100} across three LLMs. This is the average value across four subsets in PFaithBench.
Model Family
Avg. Model Size (B)
SFAcc ↑
OverAns ↓
OverAbs ↓
AnsAcc ↑
AbsAcc ↑
Leakage ↓
Llama
30.40 Billion
33.36
36.41
19.48
52.97
63.59
14.67
Qwen3 (NoThink)
10.05 Billion
42.64
33.73
14.52
64.82
66.27
10.34
Qwen3 (Think)
10.05 Billion
31.09
59.66
3.67
74.90
40.34
14.66
Qwen3.5 (NoThink)
8.56 Billion
45.47
35.28
7.06
72.05
64.72
8.80
Qwen3.5 (Think)
8.56 Billion
39.58
46.57
5.06
74.98
53.43
6.92
Qwen3.6
27.00 Billion
44.47
42.21
4.26
79.98
57.79
12.67
Appendix
Table 9: Family-level performance across base models on PFaithBench. Metrics are averaged equally across models and the four subsets. Average model sizes are in billions of parameters.
LLM
α
PFB-MultiHop
PFB-SingleHop
DFAcc ↑
SFAcc ↑
OverAns ↓
OverAbs ↓
Leakage ↓
DFAcc ↑
SFAcc ↑
OverAns ↓
OverAbs ↓
Leakage ↓
Llama-3.1-8B-Instruct
α=0
10.90
0.00
10.50
84.70
0.00
18.50
0.00
2.23
80.05
0.00
α=25
60.60
47.90
28.80
12.10
9.10
76.25
65.09
9.58
14.96
0.79
α=50
69.10
59.10
17.80
13.60
10.40
76.51
73.36
13.25
10.50
1.44
α=75
58.50
49.40
36.10
5.90
18.50
62.99
59.71
34.51
3.02
2.89
α=100
0.40
0.40
99.60
0.00
52.80
0.26
0.26
99.74
0.00
12.73
Appendix
Table 10: Detailed impact of answering-abstention data composition (asymmetric reward). The answering data percentage α vary α∈{0,25,50,75,100} , serving as a supplement to Fig. 4 .
LLM
α
PFB-MultiHop
PFB-SingleHop
DFAcc ↑
SFAcc ↑
OverAns ↓
OverAbs ↓
Leakage ↓
DFAcc ↑
SFAcc ↑
OverAns ↓
OverAbs ↓
Leakage ↓
Llama-3.1-8B-Instruct
α=0
0.00
0.00
0.00
100.00
0.00
0.00
0.00
0.00
100.00
0.00
α=25
17.00
0.30
9.30
78.90
0.00
24.67
0.52
1.97
74.67
0.00
α=50
68.10
58.70
17.10
15.40
11.90
78.35
72.97
11.68
10.24
1.44
α=75
63.90
53.60
30.00
6.30
18.90
71.39
67.59
26.38
2.76
3.02
α=100
2.70
2.10
97.30
0.10
51.50
3.67
3.54
96.33
0.00
9.32
Appendix
Table 11: Detailed impact of answering-abstention data composition (symmetric reward). The answering data percentage α vary α∈{0,25,50,75,100} , serving as a supplement to Fig. 6 .
Setting
Avg.
Prompt Strict
Inst. Strict
Prompt Loose
Inst. Loose
Setting
Prompt Strict
Inst. Strict
Prompt Loose
Inst. Loose
Llama-3.1-8B-Instruct
α=0
78.82
73.20
81.06
77.08
83.93
Base Model
76.86
70.24
79.38
74.86
82.97
α=25
77.42
71.90
79.74
75.42
82.61
MHSuff-SHInsuff
80.34
74.49
82.49
78.74
85.61
α=50
78.87
72.64
81.06
77.26
84.53
SHSuff-MHInsuff
79.53
74.31
82.01
77.26
84.53
α=75
78.53
72.46
80.70
76.89
84.05
PairContext
77.81
71.35
80.22
75.97
83.69
α=100
78.39
72.27
80.82
76.52
83.93
Llama-3.2-3B-Instruct
α=0
74.14
67.65
76.38
72.09
80.46
Appendix
Table 12: Detailed results on IFEval. IFEval reports four metrics: prompt-level strict accuracy, prompt-level loose accuracy, instruction-level strict accuracy, and instruction-level loose accuracy. In Table 5 , we report the average of these four metrics. Here, we provide the detailed results.
Data Science Institute University of Chicago Chicago, USA · Department of Computer Science University of Chicago Chicago, USA · Booth School of Business University of Chicago Chicago, USA