OpenFC: Learning Verification Policies towards Open-Search Fact Checking
Organizations: Institute of Automation, Chinese Academy of Sciences · University of Electronic Science and Technology of China · Beijing Institute of Technology · Nankai University
Abstract
Open-search fact checking is not merely retrieval followed by classification, but a sequential decision problem in which every query, source visit, and stopping decision reshapes the evidence available for verification. Yet existing systems often distribute these decisions across predefined pipelines or separately prompted modules rather than learning them as a unified task-specific policy. We introduce \textbf{OpenFC}, a unified verification-policy training framework that post-trains Qwen3-8B as a compact next-action controller over reasoning, evidence acquisition, and stopping. OpenFC learns this policy in two stages. \textbf{Stepwise-Calibrated Cold Start (SCCS)} uses a strong training-time supervisor to review post-initial reasoning, tool-use, and stopping proposals before execution, producing reliable trajectories for supervised fine-tuning without access to gold verdicts. \textbf{Verification-Aware Reinforcement Learning (VA-RL)} then improves the cold-start policy on unresolved claims through budget-aware tool rewards, label-aware advantage reweighting, and localized response masking. Across six fact-checking benchmarks, OpenFC achieves 70.39% average accuracy and 63.30% macro-F1, the highest overall averages among the evaluated methods. Stage-wise ablations further show that SCCS and VA-RL provide complementary gains, supporting the design of the two-stage training framework. These results position OpenFC as a strong and effective framework for open-search fact-checking. We will open-source our code and release the model checkpoints to support reproducibility.
Figures & tables
| Label | AVeriTeC | QuanTemp++ | SciFact | Total |
|---|---|---|---|---|
| Refuted | 1,087 | 752 | 158 | 1,997 |
| Supported | 536 | 631 | 324 | 1,491 |
| CE | 165 | 426 | – | 591 |
| NEI | 282 | – | – | 282 |
| Total | 2,070 | 1,809 | 482 | 4,361 |
| Label | AVeriTeC | QuanTemp++ | SciFact | Total |
| Refuted | 1,162 | 2,409 | 180 | 3,751 |
| Supported | 496 | 757 | 206 | 1,459 |
| CE | 194 | 1,523 | – | 1,717 |
| NEI | 556 | – | – | 556 |
| Total | 2,408 | 4,689 | 386 | 7,483 |
| Pass@3 | 1,587 | 2,894 | 313 | 4,794 |
| Method | In-distribution Dataset | Out-of-distribution Dataset | Average | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AVeriTeC | SciFact | QuanTemp++ | ClaimPlus | Climate-Fever | HealthFC | |||||||||
| Acc | F1 | Acc | F1 | Acc | F1 | Acc | F1 | Acc | F1 | Acc | F1 | Acc | F1 | |
| End-to-end Generation Models | ||||||||||||||
| DeepSeek-V4-Flash | 58.60 | 40.04 | 76.06 | 67.72 | 49.35 | 40.21 | 49.38 | 38.30 | 45.99 | 32.40 | 44.53 | 46.70 | 53.99 | 44.23 |
| GPT-5.4 | 53.40 | 36.48 | 79.26 | 76.21 | 44.07 | 37.07 | 51.25 | 41.05 | 48.60 | 34.30 | 61.47 | 55.90 | 56.34 | 46.84 |
| Claude-Sonnet-4.6 | 61.40 | 42.52 | 69.15 | 59.20 | 49.03 | 43.17 | 45.63 | 39.50 | 46.19 | 31.70 | 29.47 | 30.20 | 50.15 | 41.05 |
| Method | Avg Acc. | Tools | Visit | Search | Auto |
|---|---|---|---|---|---|
| Qwen3-8B + Tools | 50.60 | 2.28 | 1.10 | 1.18 | – |
| ClaimCheck | 49.74 | 11.72 | – | 3.08 | 8.65 |
| DEFAME | 53.43 | 14.91 | – | 4.27 | 10.64 |
| Tongyi-DR-30B | 52.35 | 8.90 | 4.24 | 4.67 | – |
| MiroThink-30B | 60.47 | 17.47 | 5.03 | 12.44 | – |
| OpenFC (Cold Start) | 57.73 | 9.92 | 2.72 | 7.20 | – |
| Method | AVeriTeC | SciFact | QuanTemp++ | ClaimPlus | Tool Usage | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Acc. | F1 | Acc. | F1 | Acc. | F1 | Acc. | F1 | Tools | Search | Visit | |
| Vanilla SFT | 65.80 | 43.14 | 63.30 | 62.55 | 61.51 | 55.53 | 60.00 | 48.95 | 9.89 | 5.08 | 4.81 |
| Stepwise-Calibration SFT | 66.20 | 46.49 | 69.15 | 65.67 | 64.49 | 57.97 | 62.50 | 53.82 | 9.92 | 7.20 | 2.72 |
| Vanilla GRPO | 71.60 | 37.80 | 81.38 | 71.75 | 73.61 | 68.90 | 52.50 | 35.84 | 8.02 | 5.99 | 2.02 |
| + Advantage reweighting | 74.00 | 41.37 | 88.83 | 83.31 | 86.28 | 82.91 | 58.75 | 40.88 | 6.43 | 4.33 | 2.11 |
| + Tool reward | 70.60 | 40.68 | 88.83 | 82.03 | 85.73 | 82.57 | 60.00 | 45.13 | 8.20 | 7.61 | 0.59 |
| Method | E-Rep. | S-Rep@0.95 | Info Gain |
|---|---|---|---|
| OpenFC | 0.19% | 3.33% | 14.67% |
| OpenFC (Cold Start) | 8.27% | 11.24% | 13.68% |
| Qwen3-8B w/ Tools | 1.08% | 1.42% | 12.79% |
| Tongyi-DR | 0.72% | 4.19% | 15.58% |
| MiroThink | 1.12% | 9.40% | 10.73% |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | SUP | REF | NEI | CON | Overall | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| P | R | F1 | P | R | F1 | P | R | F1 | P | R | F1 | Acc | F1 | |
| End-to-end Generation Models | ||||||||||||||
| DeepSeek-V4-Flash | 47.87 | 81.15 | 60.22 | 80.54 | 60.65 | 69.19 | 16.49 | 14.29 | 15.31 | 28.96 | 10.53 | 15.44 | 58.60 | 40.04 |
| GPT-5.4 | 49.20 | 85.20 | 62.38 | 86.27 | 50.45 | 63.67 | 10.36 | 23.08 | 14.30 | 96.03 | 2.87 | 5.57 | 53.40 | 36.48 |
| Claude-Sonnet-4.6 | 56.37 | 65.68 | 60.67 | 82.73 | 69.46 | 75.52 | 13.91 | 25.74 | 18.06 | 15.85 | 15.82 | 15.83 | 61.40 | 42.52 |
| Qwen3-8B | 58.14 | 35.36 | 43.97 | 70.10 | 57.98 | 63.47 | 9.43 | 37.17 | 15.04 | 20.64 | 18.45 | 19.48 | 48.00 | 35.49 |
| Method | SUP | REF | Overall | |||||
|---|---|---|---|---|---|---|---|---|
| P | R | F1 | P | R | F1 | Acc | F1 | |
| End-to-end Generation Models | ||||||||
| DeepSeek-V4-Flash | 53.44 | 80.90 | 64.36 | 70.94 | 71.22 | 71.08 | 76.06 | 67.72 |
| GPT-5.4 | 71.62 | 84.10 | 77.36 | 75.71 | 74.42 | 75.06 | 79.26 | 76.21 |
| Claude-Sonnet-4.6 | 49.77 | 64.32 | 56.12 | 53.78 | 73.98 | 62.28 | 69.15 | 59.20 |
| Qwen3-8B | 62.56 | 51.57 | 56.54 | 32.23 | 54.76 | 40.58 | 52.66 | 48.56 |
| Method | SUP | REF | CON | Overall | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| P | R | F1 | P | R | F1 | P | R | F1 | Acc | F1 | |
| End-to-end Generation Models | |||||||||||
| DeepSeek-V4-Flash | 33.56 | 78.33 | 46.99 | 78.08 | 56.12 | 65.30 | 39.88 | 4.66 | 8.34 | 49.35 | 40.21 |
| GPT-5.4 | 35.54 | 82.21 | 49.63 | 81.80 | 47.12 | 59.80 | 82.10 | 0.90 | 1.78 | 44.07 | 37.07 |
| Claude-Sonnet-4.6 | 34.99 | 64.46 | 45.36 | 77.76 | 57.96 | 66.42 | 45.74 | 11.00 | 17.73 | 49.03 | 43.17 |
| Qwen3-8B | 33.73 | 23.06 | 27.39 | 66.04 | 49.44 | 56.55 | 30.88 | 9.72 | 14.79 | 35.51 | 32.91 |
| Method | SUP | REF | NEI | CON | Overall | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| P | R | F1 | P | R | F1 | P | R | F1 | P | R | F1 | Acc | F1 | |
| End-to-end Generation Models | ||||||||||||||
| DeepSeek-V4-Flash | 55.35 | 77.09 | 64.44 | 53.86 | 66.68 | 59.59 | 14.41 | 15.38 | 14.88 | 36.46 | 8.89 | 14.29 | 49.38 | 38.30 |
| GPT-5.4 | 62.46 | 86.71 | 72.61 | 57.53 | 60.72 | 59.08 | 18.13 | 46.50 | 26.09 | 49.45 | 3.43 | 6.42 | 51.25 | 41.05 |
| Claude-Sonnet-4.6 | 60.01 | 64.58 | 62.21 | 51.21 | 57.41 | 54.13 | 13.80 | 30.77 | 19.05 | 41.37 | 15.56 | 22.61 | 45.63 | 39.50 |
| Qwen3-8B | 76.49 | 27.08 | 40.00 | 48.48 | 59.26 | 53.33 | 6.25 | 30.77 | 10.39 | 38.45 | 11.11 | 17.24 | 33.75 | 30.24 |
| Method | SUP | REF | NEI | CON | Overall | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| P | R | F1 | P | R | F1 | P | R | F1 | P | R | F1 | Acc | F1 | |
| End-to-end Generation Models | ||||||||||||||
| DeepSeek-V4-Flash | 51.27 | 78.41 | 62.00 | 32.38 | 51.03 | 39.62 | 43.89 | 11.16 | 17.80 | 17.16 | 7.24 | 10.18 | 45.99 | 32.40 |
| GPT-5.4 | 54.95 | 77.98 | 64.47 | 37.20 | 56.52 | 44.87 | 42.22 | 19.41 | 26.59 | 29.52 | 0.65 | 1.27 | 48.60 | 34.30 |
| Claude-Sonnet-4.6 | 53.47 | 81.50 | 64.57 | 33.95 | 50.59 | 40.63 | 58.02 | 8.23 | 14.42 | 9.31 | 5.84 | 7.18 | 46.19 | 31.70 |
| Qwen3-8B | 61.94 | 57.82 | 59.81 | 34.01 | 64.50 | 44.54 | 56.05 | 21.44 | 31.02 | 11.73 | 20.17 | 14.83 | 43.91 | 37.55 |
| Method | SUP | REF | NEI | Overall | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| P | R | F1 | P | R | F1 | P | R | F1 | Acc | F1 | |
| End-to-end Generation Models | |||||||||||
| DeepSeek-V4-Flash | 45.63 | 77.57 | 57.46 | 40.96 | 33.81 | 37.04 | 79.77 | 31.92 | 45.60 | 44.53 | 46.70 |
| GPT-5.4 | 49.13 | 84.16 | 62.04 | 52.11 | 29.60 | 37.75 | 78.13 | 60.05 | 67.91 | 61.47 | 55.90 |
| Claude-Sonnet-4.6 | 36.04 | 83.26 | 50.30 | 38.16 | 31.25 | 34.36 | 33.08 | 3.26 | 5.94 | 29.47 | 30.20 |
| Qwen3-8B | 52.84 | 59.90 | 56.15 | 33.98 | 28.00 | 30.70 | 78.77 | 34.98 | 48.45 | 40.53 | 45.10 |
| Configuration | Value |
|---|---|
| RL algorithm | |
| Algorithm | GRPO |
| Loss aggregation | Token mean |
| Rollouts per prompt | 8 |
| Advantage normalization | Group mean / std |
| Dynamic sampling | Enabled |
| Configuration | Value |
|---|---|
| Decoding | |
| Inference engine | vLLM |
| Temperature | |
| Top- | |
| Presence penalty | |
| Max tokens per turn | 6,000 |