cs.AISep 29, 2026

Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents

Authors: Jianchang Su, Yiwei Yang, Wei Zhang

Organizations: University of Connecticut · University of California, Santa Cruz

Abstract

Retrieval-augmented fact-checkers often receive a reliability label, such as HIGH or LOW trust, for each evidence source. These labels should adjust the model's confidence and its decision to search for more evidence, while the verdict should follow the evidence content. We introduce TrustSwap, a counterfactual test that swaps, lowers, or removes source labels while keeping every evidence text fixed, and measures its three output channels (the verdict, the confidence, and the search decision) separately. Across untrained and RL-trained models at two scales, three datasets, and two prompts, confidence and search respond to the labels as intended in 49 of 50 comparisons, yet a label change alone alters 4-23% of confident verdicts for Qwen3 models and up to 50% for an existing RL-trained fact-checker. Standard GRPO fine-tuning amplifies this shortcut at 8B in all six settings. To reduce it, we propose trust-swap augmentation (TSA), which trains GRPO on each claim with both its original and its label-swapped evidence under the same gold verdict. At 4B, TSA lowers the verdict flip rate by 7-35% (relative) in four of six settings, keeps accuracy and the intended confidence and search responses, outperforms reward-based alternatives in the main setting, and carries over to an unseen label-removal perturbation. An added consistency reward helps on the trained-on swap but not on unseen perturbations. At 8B, TSA's effect is not detectable, which makes scale the main open question.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 27, 2026cs.AI

OpenFC: Learning Verification Policies towards Open-Search Fact Checking

Open-search fact checking is not merely retrieval followed by classification, but a sequential decision problem in which every query, source visit, and stopping decision reshapes the evidence available for verification. Yet existing systems often distribute these decisions across predefined pipelines or separately prompted modules rather than learning them as a unified task-specific policy. We introduce \textbf{OpenFC}, a unified verification-policy training framework that post-trains Qwen3-8B as a compact next-action controller over reasoning, evidence acquisition, and stopping. OpenFC learns this policy in two stages. \textbf{Stepwise-Calibrated Cold Start (SCCS)} uses a strong training-time supervisor to review post-initial reasoning, tool-use, and stopping proposals before execution, producing reliable trajectories for supervised fine-tuning without access to gold verdicts. \textbf{Verification-Aware Reinforcement Learning (VA-RL)} then improves the cold-start policy on unresolved claims through budget-aware tool rewards, label-aware advantage reweighting, and localized response masking. Across six fact-checking benchmarks, OpenFC achieves 70.39% average accuracy and 63.30% macro-F1, the highest overall averages among the evaluated methods. Stage-wise ablations further show that SCCS and VA-RL provide complementary gains, supporting the design of the two-stage training framework. These results position OpenFC as a strong and effective framework for open-search fact-checking. We will open-source our code and release the model checkpoints to support reproducibility.
Sep 14, 2026cs.CL

Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation

Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuality-based metrics, their sensitivity and reliability remain underexplored. This paper introduces a meta-evaluation framework that systematically tests these metrics using controlled corruptions of gold standard answers. Our method generates ranked outputs with known degrees of degradation to probe how metrics capture nuanced changes in truthfulness. Our experiments reveal that pipeline-based methods, such as the RAGAS's factual correctness metric, better track degradation than LLM-as-judge approaches. We also propose a new variant of the factual correctness metric that provides a competitive and cost-efficient.
Dec 29, 2025cs.HC

Althea: The Fact-Checking--Metalearning Tradeoff in AI-Assisted Verification

Fact-checking systems must be scalable and epistemically trustworthy. We introduce Althea, a retrieval-augmented system for user-driven claim evaluation that matches standard pipelines on AVeriTeC while improving supported/refuted discrimination. A longitudinal survey experiment (N=961) treats a ten-day follow-up as a fading test: after modeling a verification procedure, we remove the system and ask whether users reproduce it unaided, testing metalearning rather than one-time accuracy. We compare two AI-assisted treatments, Exploratory (guided reasoning) and Summary (synthesized verdicts), against two baselines, unrelated news and Self-search. The treatments yield the strongest immediate accuracy and confidence gains but do not survive the fading test: on unseen claims they perform no better than news, while Self-search, with no procedure to fade, retains a large advantage. This reveals a factchecking-metalearning tradeoff: conditions that most improve immediate accuracy are least likely to produce metalearning, cautioning against treating AI-delivered verdicts as a source of durable literacy gains.