cs.CLOct 5, 2026

SAFE-MR: Evidence Sufficiency Learning for Selective Multimodal Rumor Detection

Authors: Shiwen Ni

Organizations: Artificial Intelligence Research Institute, Shenzhen University of Advanced Technology

Abstract

Multimodal rumor detectors increasingly rely on retrieved evidence, yet relevant evidence is not necessarily sufficient for verification. Missing provenance, duplicated reports, and unresolved contradictions can produce confident predictions without adequate support. We introduce SAFE-MR, a framework that separates claim veracity from evidence sufficiency. The method decomposes image-text posts into verifiable claims, constructs a relation-aware claim-evidence graph, and aggregates evidence using provenance and contextual compatibility. Separate veracity and sufficiency heads support selective prediction, while evidence interventions encourage stability under irrelevant additions and sensitivity to evidence removal. On NewsCLIPpings, VERITE, and XFacta, SAFE-MR achieves macro-F1 scores of 91.2%, 75.8%, and 85.2%, respectively. Against the matched backbone with evidence, its macro-F1 gains are 2.2, 4.9, and 4.8 percentage points. On the diagnostic selection set, SAFE-MR reduces AURC from 0.105 for maximum-probability rejection to 0.075 and lowers error at 80% coverage from 13.8% to 8.5%. Evidence-perturbation and ablation results support the role of sufficiency learning and intervention training in improving selective verification.

Figures & tables

Explore similar work

CardsList
  1. Multimodal rumor detection enhanced by external evidence and forgery features

    Jan 21, 2026Han Li, Hua SunMultimodal FusionForgeries

  2. RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild

    May 11, 2026Danni Xu, Shaojing Fan, Harry Cheng +1Fact-CheckingReference-Guided Multimodal In-Context Verification

  3. Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection

    Jul 20, 2026Haochen Zhao, Yongxiu Xu, Xinkui Lin +6Misinformation DetectionFine-Grained Video Understanding