cs.AISep 29, 2026

Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents

Authors: Jianchang Su, Yiwei Yang, Wei Zhang

Organizations: University of Connecticut · University of California, Santa Cruz

Abstract

Retrieval-augmented fact-checkers often receive a reliability label, such as HIGH or LOW trust, for each evidence source. These labels should adjust the model's confidence and its decision to search for more evidence, while the verdict should follow the evidence content. We introduce TrustSwap, a counterfactual test that swaps, lowers, or removes source labels while keeping every evidence text fixed, and measures its three output channels (the verdict, the confidence, and the search decision) separately. Across untrained and RL-trained models at two scales, three datasets, and two prompts, confidence and search respond to the labels as intended in 49 of 50 comparisons, yet a label change alone alters 4-23% of confident verdicts for Qwen3 models and up to 50% for an existing RL-trained fact-checker. Standard GRPO fine-tuning amplifies this shortcut at 8B in all six settings. To reduce it, we propose trust-swap augmentation (TSA), which trains GRPO on each claim with both its original and its label-swapped evidence under the same gold verdict. At 4B, TSA lowers the verdict flip rate by 7-35% (relative) in four of six settings, keeps accuracy and the intended confidence and search responses, outperforms reward-based alternatives in the main setting, and carries over to an unseen label-removal perturbation. An added consistency reward helps on the trained-on swap but not on unseen perturbations. At 8B, TSA's effect is not detectable, which makes scale the main open question.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. OpenFC: Learning Verification Policies towards Open-Search Fact Checking

    Sep 27, 2026Xinming Wang, Kaixiang Qiu, Yansong Lin +5Fact-Checking

  2. Can We Trust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation

    Sep 14, 2026Sarra Gharsallah, Adele Robaldo, Mariia Tokareva +5JudgesValidation

  3. Althea: The Fact-Checking--Metalearning Tradeoff in AI-Assisted Verification

    Dec 29, 2025Svetlana Churina, Kokil Jaidka, Anab Maulana Barik +5Fact-CheckingCategory-Aware Atomic Claims