Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipulated material into complex forms in which native and generated content jointly participate. Existing multimodal fake news detection research primarily focuses on veracity assessment and rarely characterizes how generativity differences affect the reliability of evidence. In contrast, AIGC detection primarily determines whether content is generated or modified by generative models, but it does not by itself establish whether the underlying news event is true. To bridge the separation between these tasks in data and evaluation, we construct Weibo26, a multimodal fake news detection dataset for generative-content scenarios. On this basis, we propose the Generativity-Aware Hierarchical Reasoning (GAHR) framework, which combines global judgment with local correction so that generativity information participates in news-veracity reasoning. Experiments on multiple existing fake news detection benchmarks and Weibo26 show that GAHR achieves competitive veracity-detection performance while effectively identifying generative content.
Figures & tables
Figure 1: Conceptual organization of multimodal news samples along two distinct dimensions: veracity (real versus fake) and generative involvement (native versus generated). Native samples contain no generated content, whereas generated-content samples contain AI-generated content.
Figure 2: Overview of the six sample types in Weibo26, including their statistics and distribution.
Dataset
Language
Publication year
Time range
Real
Fake
Image
PHEME ( Kochkina et al., 2018 )
English
2017
2014–2015
3,830
1,972
5,802
GossipCop ( Shu et al., 2020 )
English
2020
up to 2020
16,817
5,323
18,417
Fakeddit ( Nakamura et al., 2020 )
English
2020
2008–2019
527,049
628,501
682,996
AMG ( Guo et al., 2025 )
English
2025
2016–2024
3,018
2,004
5,022
Weibo17 ( Jin et al., 2017 )
Chinese
2017
2012–2016
4,779
4,749
9,528
Weibo21 ( Qian et al., 2021 )
Chinese
2021
2014–2021
4,640
4,488
9,128
Table 1: Comparison of Weibo26 with existing multimodal fake news detection datasets.
Figure 3: Overall architecture of the Generativity-Aware Hierarchical Reasoning (GAHR) framework.
Dataset
Method
Year
Accuracy
Fake News
Real News
AUC
Precision
Recall
F1-score
Precision
Recall
F1-score
PHEME
MVAE
2019
0.774 ± 0.007
0.677 ± 0.007
0.565 ± 0.006
0.616 ± 0.006
0.810 ± 0.006
0.858 ± 0.006
0.833 ± 0.006
0.841 ± 0.007
HMCAN
2021
0.840 ± 0.014
0.724 ± 0.014
0.736 ± 0.014
0.730 ± 0.014
0.889 ± 0.015
0.883 ± 0.015
0.886 ± 0.014
0.907 ± 0.014
CAFE
2022
0.816 ± 0.012
0.724 ± 0.012
0.606 ± 0.013
0.659 ± 0.013
0.847 ± 0.012
0.904 ± 0.013
0.873 ± 0.013
0.884 ± 0.012
MRML
2023
0.853 ± 0.007
0.722 ± 0.007
0.811 ± 0.007
0.764 ± 0.007
0.917 ± 0.008
0.870 ± 0.007
0.893 ± 0.007
0.924 ± 0.006
NSLM
2024
0.826 ± 0.006
0.715 ± 0.005
0.679 ± 0.006
0.697 ± 0.006
0.869 ± 0.006
0.887 ± 0.007
0.878 ± 0.006
0.897 ± 0.006
Table 2: Veracity detection results on six datasets. The results are reported as mean ± standard deviation over five runs. Red and blue indicate the best and second-best results, respectively.
Dataset
Method
Accuracy
Fake News
Real News
AUC
Precision
Recall
F1-score
Precision
Recall
F1-score
AMG
w/o Text
0.778 ± 0.010
0.622 ± 0.010
0.773 ± 0.010
0.689 ± 0.009
0.880 ± 0.010
0.781 ± 0.009
0.827 ± 0.009
0.830 ± 0.010
w/o Image
0.809 ± 0.015
0.697 ± 0.014
0.797 ± 0.015
0.744 ± 0.015
0.873 ± 0.014
0.816 ± 0.014
0.842 ± 0.015
0.866 ± 0.015
w/o Global
0.817 ± 0.007
0.708 ± 0.007
0.757 ± 0.007
0.732 ± 0.006
0.872 ± 0.007
0.845 ± 0.006
0.858 ± 0.007
0.894 ± 0.006
w/o Local
0.817 ± 0.014
0.679 ± 0.014
0.844 ± 0.014
0.754 ± 0.014
0.908 ± 0.014
0.806 ± 0.013
0.854 ± 0.014
0.891 ± 0.013
Full Model
0.843 ± 0.007
0.706 ± 0.007
0.870 ± 0.006
0.779 ± 0.006
0.932 ± 0.007
0.830 ± 0.007
0.878 ± 0.007
0.925 ± 0.006
Table 3: Ablation results of GAHR on AMG, Weibo21, and Weibo26. The results are reported as mean ± standard deviation over five runs. Red and blue indicate the best and second-best results, respectively.
Method
Accuracy
Generated Content
Native Content
AUC
Precision
Recall
F1-score
Precision
Recall
F1-score
BERT
0.770 ± 0.013
0.681 ± 0.012
0.638 ± 0.013
0.659 ± 0.012
0.810 ± 0.013
0.838 ± 0.013
0.824 ± 0.013
0.844 ± 0.013
RoBERTa
0.780 ± 0.008
0.696 ± 0.007
0.655 ± 0.007
0.674 ± 0.008
0.820 ± 0.007
0.845 ± 0.008
0.832 ± 0.007
0.855 ± 0.008
MPU
0.800 ± 0.014
0.727 ± 0.014
0.678 ± 0.014
0.702 ± 0.014
0.832 ± 0.015
0.863 ± 0.015
0.847 ± 0.015
0.877 ± 0.015
DP-Net
0.810 ± 0.013
0.746 ± 0.013
0.688 ± 0.013
0.716 ± 0.013
0.838 ± 0.013
0.873 ± 0.014
0.855 ± 0.014
0.890 ± 0.013
GAHR-T (Ours)
0.821 ± 0.013
0.766 ± 0.013
0.704 ± 0.012
0.733 ± 0.013
0.848 ± 0.013
0.883 ± 0.013
0.865 ± 0.014
0.908 ± 0.012
Table 4: Unimodal text-side AIGC detection results on Weibo26. Red and blue indicate the best and second-best results, respectively.
Method
Accuracy
Generated Content
Native Content
AUC
Precision
Recall
F1-score
Precision
Recall
F1-score
LGrad
0.822 ± 0.006
0.676 ± 0.006
0.759 ± 0.006
0.715 ± 0.006
0.892 ± 0.006
0.845 ± 0.006
0.867 ± 0.006
0.866 ± 0.005
FreqNet
0.833 ± 0.011
0.692 ± 0.012
0.773 ± 0.011
0.730 ± 0.011
0.897 ± 0.011
0.851 ± 0.012
0.875 ± 0.011
0.876 ± 0.011
SAFE
0.882 ± 0.007
0.776 ± 0.007
0.839 ± 0.007
0.806 ± 0.006
0.928 ± 0.006
0.896 ± 0.006
0.912 ± 0.007
0.924 ± 0.006
UnivFD
0.935 ± 0.015
0.860 ± 0.014
0.908 ± 0.015
0.884 ± 0.015
0.958 ± 0.014
0.937 ± 0.016
0.947 ± 0.014
0.966 ± 0.015
DFFreq
0.922 ± 0.010
0.841 ± 0.010
0.899 ± 0.010
0.869 ± 0.010
0.955 ± 0.010
0.927 ± 0.010
0.941 ± 0.010
0.957 ± 0.010
Table 5: Unimodal image-side AIGC detection results on Weibo26. Red and blue indicate the best and second-best results, respectively.
Figure 4: Veracity detection with and without generativity labels across the six Weibo26 sample types.
Figure 5: t-SNE visualization of generativity-related representations before and after GAHR processing.
Figure 6: Effects of the attention temperature τ and the generativity-detection loss weight λG on model performance.
Figure 7: Representative cases under different types of generative involvement.
The rapid dissemination of multimodal content has intensified the spread of fabricated news, presenting a substantial threat to social integrity. A formidable challenge for current detection systems is identifying misinformation related to novel events in zero-shot scenarios. Prevailing zero-shot methods typically assess news items in isolation via semantic matching, a strategy that fails to recognize the recycled disinformation tactics from past campaigns and lacks the sophisticated reasoning needed to identify subtle, cross-modal discrepancies. To surmount these deficiencies, we introduce \textbf{MRAFnd}, a novel \underline{\textbf{M}}ultimodal \underline{\textbf{R}}etrieval-\underline{\textbf{A}}ugmented Framework for Zero-Shot \underline{\textbf{F}}ake \underline{\textbf{N}}ews \underline{\textbf{D}}etection. MRAFnd emulates a collaborative team of analysts to verify news veracity. The framework initiates with \textbf{Multimodal Similarity-based News Retrieval} to assemble a corpus of contextually analogous articles from an unlabeled reference database. Subsequently, during the \textbf{Bifurcated Evidential Reasoning} stage, agents perform a dual-directional analysis to extract critical patterns from the retrieved evidence. Finally, a \textbf{Multi-Agent Collaborative Debate}, involving Analyst and Arbiter agents, engages in a structured discourse to arrive at a definitive and robust conclusion. Comprehensive experiments on three benchmark datasets reveal that MRAFnd markedly surpasses state-of-the-art baselines, achieving an accuracy gain of up to 2.35% on the demanding Weibo-21 dataset.
Lehan Zhang, Yinlei Cheng, Shiqi Hu Yiheng Zhou +2
Beijing Institute of Fashion Technology, Beijing, China · Shenyang University of Technology, Shenyang, China
In recent years, multimodal multidomain fake news detection has garnered increasing attention. Nevertheless, this direction presents two significant challenges: (1) Failure to Capture Cross-Instance Narrative Consistency: existing models usually evaluate each news in isolation, fail to capture cross-instance narrative consistency, and thus struggle to address the spread of cluster based fake news driven by social media; (2) Lack of Domain Specific Knowledge for Reasoning: conventional models, which rely solely on knowledge encoded in their parameters during training, struggle to generalize to new or data-scarce domains (e.g., emerging events or niche topics). To tackle these challenges, we introduce Retrieval-Augmented Multimodal Model for Fake News Detection (RAMM). First, RAMM employs a Multimodal Large Language Model (MLLM) as its backbone to capture cross-modal semantic information from news samples. Second, RAMM incorporates an Abstract Narrative Alignment Module. This component adaptively extracts abstract narrative consistency from diverse instances across distinct domains, aggregates relevant knowledge, and thereby enables the modeling of high-level narrative information. Finally, RAMM introduces a Semantic Representation Alignment Module, which aligns the model's decision-making paradigm with that of humans - specifically, it shifts the model's reasoning process from direct inference on multimodal features to an instance-based analogical reasoning process. Extensive experimental results on three public datasets validate the efficacy of our proposed approach. Our code is available at the following link: https://github.com/li-yiheng/RAMM
Yiheng Li, Weihai Lu, Hanyi Yu +1
University of International Business and Economics · Beijing, China · Peking University +4
Existing multimodal fake news detection methods often introduce external information to assist detection. However, most of them rely on entity-level retrieval and are therefore prone to introducing event-irrelevant noise. Meanwhile, existing methods mainly focus on improving overall performance and do not account for differences in the degree of harm posed by different instances of fake news. To address these limitations, we design an Event-Level Evidence Retrieval Framework (ELERF) and propose a Relation-Aware Evidence Graph Network (RAEGNet). ELERF retrieves external evidence based on the complete event semantics of a news item. RAEGNet constructs a directed graph that incorporates news-evidence stance relations and evidence-evidence interaction relations, and introduces a conditional-harm branch to jointly model authenticity and potential harm. Experimental results demonstrate that RAEGNet outperforms multiple baseline methods across all evaluated metrics on Weibo-21, Fakeddit, and our self-constructed SSS dataset.
Wenbin Shen, Guoxuan Qin, Guangxu Yao +4
Nanjing University of Science and Technology · Zhejiang University · University of Chinese Academy of Sciences