cs.CLSep 30, 2026

Towards Robust Numerical Claim Verification

Authors: Peter Røysland Aarnes, Vinay Setty

Organizations: University of Stavanger · Factiverse AI and University of Stavanger

Abstract

Large language models (LLMs) are widely used for claim verification, yet remain brittle for numerical reasoning: even small changes in value can sharply degrade accuracy. We show that this brittleness persists in frontier LLMs, but can be mitigated through adversarial fine-tuning on numerically perturbed examples. Using parameter-efficient fine-tuning, small Qwen3 models (0.6B\unicodex2013\unicode{x2013}8B) reach 98.7% accuracy on label-flipping perturbations, outperforming larger zero-shot models and frontier systems (GPT-5.4 Pro (74.0%) and Gemini 2.5 Flash (73.9%)). The gains generalise to unseen perturbation types, indicating robust numerical decision boundaries rather than memorised edits. Robustness also transfers without target-domain data, significantly improving cross-lingual performance in Spanish. We further show that the same fine-tuning recipe confers robustness to evidence-side perturbations, using the VitaminC dataset.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification

    Jul 27, 2026Sagnik Sinha, Shreyas ShresthaCategory-Aware Atomic ClaimsLLM Reasoning Strategies

  2. Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks

    Jun 2, 2026Malia Barker, Bishal Lakha, Edoardo Serra +1Mathematical Reasoning BenchmarksLLM Reasoning Strategies