Organizations: Institute of Automation, Chinese Academy of Sciences · University of Electronic Science and Technology of China · Beijing Institute of Technology · Nankai University
Open-search fact checking is not merely retrieval followed by classification, but a sequential decision problem in which every query, source visit, and stopping decision reshapes the evidence available for verification. Yet existing systems often distribute these decisions across predefined pipelines or separately prompted modules rather than learning them as a unified task-specific policy. We introduce \textbf{OpenFC}, a unified verification-policy training framework that post-trains Qwen3-8B as a compact next-action controller over reasoning, evidence acquisition, and stopping. OpenFC learns this policy in two stages. \textbf{Stepwise-Calibrated Cold Start (SCCS)} uses a strong training-time supervisor to review post-initial reasoning, tool-use, and stopping proposals before execution, producing reliable trajectories for supervised fine-tuning without access to gold verdicts. \textbf{Verification-Aware Reinforcement Learning (VA-RL)} then improves the cold-start policy on unresolved claims through budget-aware tool rewards, label-aware advantage reweighting, and localized response masking. Across six fact-checking benchmarks, OpenFC achieves 70.39% average accuracy and 63.30% macro-F1, the highest overall averages among the evaluated methods. Stage-wise ablations further show that SCCS and VA-RL provide complementary gains, supporting the design of the two-stage training framework. These results position OpenFC as a strong and effective framework for open-search fact-checking. We will open-source our code and release the model checkpoints to support reproducibility.
Figures & tables
Figure 1: Pipeline-centric automatic fact-checking (AFC) assigns verification to staged modules, whereas agentic AFC integrates reasoning, retrieval, and stopping in a ReAct-style.
Figure 2: OpenFC training process. Stage 1 uses Stepwise-Calibrated Cold Start (SCCS) to calibrate policy rollouts before execution; Stage 2 applies Verification-Aware Reinforcement Learning (VA-RL) to improve fact-checking generalization.
Label
AVeriTeC
QuanTemp++
SciFact
Total
Refuted
1,087
752
158
1,997
Supported
536
631
324
1,491
CE
165
426
–
591
NEI
282
–
–
282
Total
2,070
1,809
482
4,361
Table 1: Accepted claim-level trajectories for cold-start SFT.
Label
AVeriTeC
QuanTemp++
SciFact
Total
Refuted
1,162
2,409
180
3,751
Supported
496
757
206
1,459
CE
194
1,523
–
1,717
NEI
556
–
–
556
Total
2,408
4,689
386
7,483
Pass@3
1,587
2,894
313
4,794
Table 2: Label distribution of the RL training data.
Method
In-distribution Dataset
Out-of-distribution Dataset
Average
AVeriTeC
SciFact
QuanTemp++
ClaimPlus
Climate-Fever
HealthFC
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
F1
Acc
F1
End-to-end Generation Models
DeepSeek-V4-Flash
58.60
40.04
76.06
67.72
49.35
40.21
49.38
38.30
45.99
32.40
44.53
46.70
53.99
44.23
GPT-5.4
53.40
36.48
79.26
76.21
44.07
37.07
51.25
41.05
48.60
34.30
61.47
55.90
56.34
46.84
Claude-Sonnet-4.6
61.40
42.52
69.15
59.20
49.03
43.17
45.63
39.50
46.19
31.70
29.47
30.20
50.15
41.05
Table 3: Main results on six fact-checking datasets, including three ID and three OOD benchmarks. We report accuracy and dataset-specific macro-F1. Bold and underline indicate the best and second-best results, respectively.
Method
Avg Acc.
Tools
Visit
Search
Auto
Qwen3-8B + Tools
50.60
2.28
1.10
1.18
–
ClaimCheck
49.74
11.72
–
3.08
8.65
DEFAME
53.43
14.91
–
4.27
10.64
Tongyi-DR-30B
52.35
8.90
4.24
4.67
–
MiroThink-30B
60.47
17.47
5.03
12.44
–
OpenFC (Cold Start)
57.73
9.92
2.72
7.20
–
Table 4: Average tool use per sample. Tools is the total number of explicit Search and Visit calls; Auto reports automatic page fetching for ClaimCheck and DEFAME.
Method
AVeriTeC
SciFact
QuanTemp++
ClaimPlus
Tool Usage
Acc.
F1
Acc.
F1
Acc.
F1
Acc.
F1
Tools
Search
Visit
Vanilla SFT
65.80
43.14
63.30
62.55
61.51
55.53
60.00
48.95
9.89
5.08
4.81
Stepwise-Calibration SFT
66.20
46.49
69.15
65.67
64.49
57.97
62.50
53.82
9.92
7.20
2.72
Vanilla GRPO
71.60
37.80
81.38
71.75
73.61
68.90
52.50
35.84
8.02
5.99
2.02
+ Advantage reweighting
74.00
41.37
88.83
83.31
86.28
82.91
58.75
40.88
6.43
4.33
2.11
+ Tool reward
70.60
40.68
88.83
82.03
85.73
82.57
60.00
45.13
8.20
7.61
0.59
Table 5: Ablation study of the proposed training components. The best and second-best task-performance results are highlighted in bold and underlined, respectively.
Method
E-Rep. ↓
S-Rep@0.95 ↓
Info Gain ↑
OpenFC
0.19%
3.33%
14.67%
OpenFC (Cold Start)
8.27%
11.24%
13.68%
Qwen3-8B w/ Tools
1.08%
1.42%
12.79%
Tongyi-DR
0.72%
4.19%
15.58%
MiroThink
1.12%
9.40%
10.73%
Table 6: Analysis on information quality across different methods on QuanTemp++.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Method
SUP
REF
NEI
CON
Overall
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Acc
F1
End-to-end Generation Models
DeepSeek-V4-Flash
47.87
81.15
60.22
80.54
60.65
69.19
16.49
14.29
15.31
28.96
10.53
15.44
58.60
40.04
GPT-5.4
49.20
85.20
62.38
86.27
50.45
63.67
10.36
23.08
14.30
96.03
2.87
5.57
53.40
36.48
Claude-Sonnet-4.6
56.37
65.68
60.67
82.73
69.46
75.52
13.91
25.74
18.06
15.85
15.82
15.83
61.40
42.52
Qwen3-8B
58.14
35.36
43.97
70.10
57.98
63.47
9.43
37.17
15.04
20.64
18.45
19.48
48.00
35.49
Appendix
Table 7: Label-level results on AVeriTeC. The labels in the dataset are Supported / Refuted / NEI / Conflicting Evidence.
Method
SUP
REF
Overall
P
R
F1
P
R
F1
Acc
F1
End-to-end Generation Models
DeepSeek-V4-Flash
53.44
80.90
64.36
70.94
71.22
71.08
76.06
67.72
GPT-5.4
71.62
84.10
77.36
75.71
74.42
75.06
79.26
76.21
Claude-Sonnet-4.6
49.77
64.32
56.12
53.78
73.98
62.28
69.15
59.20
Qwen3-8B
62.56
51.57
56.54
32.23
54.76
40.58
52.66
48.56
Appendix
Table 8: Label-level results on SciFact. The labels in the dataset are Supported / Refuted.
Method
SUP
REF
CON
Overall
P
R
F1
P
R
F1
P
R
F1
Acc
F1
End-to-end Generation Models
DeepSeek-V4-Flash
33.56
78.33
46.99
78.08
56.12
65.30
39.88
4.66
8.34
49.35
40.21
GPT-5.4
35.54
82.21
49.63
81.80
47.12
59.80
82.10
0.90
1.78
44.07
37.07
Claude-Sonnet-4.6
34.99
64.46
45.36
77.76
57.96
66.42
45.74
11.00
17.73
49.03
43.17
Qwen3-8B
33.73
23.06
27.39
66.04
49.44
56.55
30.88
9.72
14.79
35.51
32.91
Appendix
Table 9: Label-level results on QuanTemp++. The labels in the dataset are Supported / Refuted / Conflicting Evidence.
Method
SUP
REF
NEI
CON
Overall
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Acc
F1
End-to-end Generation Models
DeepSeek-V4-Flash
55.35
77.09
64.44
53.86
66.68
59.59
14.41
15.38
14.88
36.46
8.89
14.29
49.38
38.30
GPT-5.4
62.46
86.71
72.61
57.53
60.72
59.08
18.13
46.50
26.09
49.45
3.43
6.42
51.25
41.05
Claude-Sonnet-4.6
60.01
64.58
62.21
51.21
57.41
54.13
13.80
30.77
19.05
41.37
15.56
22.61
45.63
39.50
Qwen3-8B
76.49
27.08
40.00
48.48
59.26
53.33
6.25
30.77
10.39
38.45
11.11
17.24
33.75
30.24
Appendix
Table 10: Label-level results on ClaimPlus. The labels in the dataset are Supported / Refuted / NEI / Conflicting Evidence.
Method
SUP
REF
NEI
CON
Overall
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Acc
F1
End-to-end Generation Models
DeepSeek-V4-Flash
51.27
78.41
62.00
32.38
51.03
39.62
43.89
11.16
17.80
17.16
7.24
10.18
45.99
32.40
GPT-5.4
54.95
77.98
64.47
37.20
56.52
44.87
42.22
19.41
26.59
29.52
0.65
1.27
48.60
34.30
Claude-Sonnet-4.6
53.47
81.50
64.57
33.95
50.59
40.63
58.02
8.23
14.42
9.31
5.84
7.18
46.19
31.70
Qwen3-8B
61.94
57.82
59.81
34.01
64.50
44.54
56.05
21.44
31.02
11.73
20.17
14.83
43.91
37.55
Appendix
Table 11: Label-level results on Climate-Fever. The labels in the dataset are Supported / Refuted / NEI / Conflicting Evidence.
Method
SUP
REF
NEI
Overall
P
R
F1
P
R
F1
P
R
F1
Acc
F1
End-to-end Generation Models
DeepSeek-V4-Flash
45.63
77.57
57.46
40.96
33.81
37.04
79.77
31.92
45.60
44.53
46.70
GPT-5.4
49.13
84.16
62.04
52.11
29.60
37.75
78.13
60.05
67.91
61.47
55.90
Claude-Sonnet-4.6
36.04
83.26
50.30
38.16
31.25
34.36
33.08
3.26
5.94
29.47
30.20
Qwen3-8B
52.84
59.90
56.15
33.98
28.00
30.70
78.77
34.98
48.45
40.53
45.10
Appendix
Table 12: Label-level results on HealthFC. The labels in the dataset are Supported / Refuted / NEI.
Large language models are increasingly used for automated fact checking, but end-to-end prompting often entangles evidence retrieval, reasoning, and uncertainty estimation, making failures difficult to diagnose and confidence difficult to trust. We present R2VC, a modular retrieve, reason, verify, calibrate architecture for evidence-grounded fact checking with citations and abstention. R2VC combines hybrid sparse+dense retrieval over Wikipedia, a supervised fine-tuned and DPO-aligned generator that produces diverse structured verdict candidates, an external NLI cross-encoder for evidence-based candidate selection, and a lightweight sequence-level calibrator for confidence estimation and selective abstention. On FEVER, an 8B backbone with R2VC achieves 13.74% higher accuracy than baseline. Ablation studies show that verifier-based candidate selection and confidence calibration are the largest contributors to performance. Removing candidate selection drops FEVER accuracy to 76.24%, while removing calibration nearly doubles the Brier score to 0.161. A manual analysis of 250 errors further shows that retrieval failures, especially wrong-entity evidence, remain the dominant bottleneck. Together, these results show that modular fact-checking pipelines can substantially improve both predictive accuracy and confidence reliability in open-domain verification.
Dhruv Dixit, Paritosh Pandey
Department of Electrical and Computer Engineering Stevens Institute of Technology Hoboken, New Jersey, United States · Department of Computer Science University of North Carolina Chapel Hill, North Carolina, United States
Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information. At the same time, knowledge-graph-based fact-checkers deliver precise and interpretable evidence, yet suffer from limited coverage or latency. By integrating LLMs with knowledge graphs and real-time search agents, we introduce a hybrid fact-checking approach that leverages the individual strengths of each component. Our system comprises three autonomous steps: 1) a Knowledge Graph (KG) Retrieval for rapid one-hop lookups in DBpedia, 2) an LM-based classification guided by a task-specific labeling prompt, producing outputs with internal rule-based logic, and 3) a Web Search Agent invoked only when KG coverage is insufficient. Our pipeline achieves an F1 score of 0.93 on the FEVER benchmark on the Supported/Refuted split without task-specific fine-tuning. To address Not enough information cases, we conduct a targeted reannotation study showing that our approach frequently uncovers valid evidence for claims originally labeled as Not Enough Information (NEI), as confirmed by both expert annotators and LLM reviewers. With this paper, we present a modular, opensource fact-checking pipeline with fallback strategies and generalization across datasets.
Shaghayegh Kolli, Richard Rosenbaum, Timo Cavelius +3
Automated fact-checking remains a challenge for Large Language Models (LLMs) due to "query brittleness" in traditional retrieval systems. We propose DeLIVeR (Decomposed Learning for Information-grounded Veracity Recognition), a framework that treats evidence retrieval as a reinforced strategic exploration task. DeLIVeR utilizes a Planner LLM to decompose complex claims into targeted question sets, which are used to traverse structured Knowledge Graphs (KGs) for high-precision evidence. We optimize the Planner's policy using Group Relative Policy Optimization (GRPO) with a reward system prioritizing structural diversity and verdict accuracy. Our evaluation on LIAR, FEVER, and PolitiFact shows that DeLIVeR significantly outperforms state-of-the-art baselines. Using Qwen2.5-7B, our framework achieved peak F1-scores of 83.73, 84.57, and 79.70 respectively, representing a 10-15% improvement over HippoRAG2. By shifting to a reinforced question-planning strategy, DeLIVeR effectively bridges multi-hop reasoning gaps and provides an auditable, transparent path for verifiable misinformation detection.
Cong Hoan Nguyen, Thomas Hoang, Hieu Minh Duong +1
University of Louisville, Louisville KY 40292, USA. · Denison University, Granville, Ohio 43023, USA.