RAEGNet: Relation-Aware Evidence Graph Network for Harm-Aware Multimodal Fake News Detection
Organizations: Nanjing University of Science and Technology · Zhejiang University · University of Chinese Academy of Sciences
Abstract
Existing multimodal fake news detection methods often introduce external information to assist detection. However, most of them rely on entity-level retrieval and are therefore prone to introducing event-irrelevant noise. Meanwhile, existing methods mainly focus on improving overall performance and do not account for differences in the degree of harm posed by different instances of fake news. To address these limitations, we design an Event-Level Evidence Retrieval Framework (ELERF) and propose a Relation-Aware Evidence Graph Network (RAEGNet). ELERF retrieves external evidence based on the complete event semantics of a news item. RAEGNet constructs a directed graph that incorporates news-evidence stance relations and evidence-evidence interaction relations, and introduces a conditional-harm branch to jointly model authenticity and potential harm. Experimental results demonstrate that RAEGNet outperforms multiple baseline methods across all evaluated metrics on Weibo-21, Fakeddit, and our self-constructed SSS dataset.
Figures & tables
| Components | Weibo-21 | Fakeddit | SSS |
|---|---|---|---|
| News | 4,493 | 5,919 | 7,997 |
| Evidence | 8,341 | 11,375 | 15,157 |
| Category | Method | Weibo-21 | Fakeddit | SSS | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Accuracy | F1-score | Accuracy | F1-score | Accuracy | F1-score | |||||
| Fake | Real | Fake | Real | Fake | Real | |||||
| w/o EK | CAFE | 0.815 | 0.810 | 0.820 | 0.786 | 0.816 | 0.745 | 0.726 | 0.753 | 0.693 |
| MRML | 0.903 | 0.908 | 0.898 | 0.860 | 0.876 | 0.839 | 0.758 | 0.779 | 0.732 | |
| Event-Radar | 0.881 | 0.884 | 0.877 | 0.840 | 0.859 | 0.815 | 0.754 | 0.767 | 0.739 | |
| MSACA | 0.894 | 0.900 | 0.886 | 0.846 | 0.864 | 0.823 | 0.762 | 0.778 | 0.744 | |
| Ablation Settings | Weibo-21 | Fakeddit | SSS | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Accuracy | HHF | Accuracy | HHF | Accuracy | HHF | ||||
| F1 | Rec. | F1 | Rec. | F1 | Rec. | ||||
| w/o Evidence Graph | 0.913 | 0.956 | 0.961 | 0.909 | 0.939 | 0.952 | 0.822 | 0.892 | 0.956 |
| – w/o Edge Weights | 0.920 | 0.959 | 0.956 | 0.915 | 0.936 | 0.935 | 0.837 | 0.886 | 0.923 |
| – w/o Edge Types | 0.914 | 0.960 | 0.967 | 0.911 | 0.930 | 0.935 | 0.826 | 0.877 | 0.928 |
| w/o Harm-aware Branch | 0.915 | 0.946 | 0.919 | 0.912 | 0.925 | 0.907 | 0.839 | 0.878 | 0.905 |
| Dataset | Retrieval | Acc. | Fake F1 | Real F1 |
|---|---|---|---|---|
| Weibo-21 | Entity-level | 0.911 | 0.911 | 0.911 |
| Event-level | 0.931 | 0.934 | 0.929 | |
| Fakeddit | Entity-level | 0.898 | 0.908 | 0.886 |
| Event-level | 0.919 | 0.926 | 0.910 | |
| SSS | Entity-level | 0.808 | 0.817 | 0.798 |
| Event-level | 0.843 | 0.860 | 0.822 |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Role | Content |
|---|---|
| System | You are an expert in fake news research data screening, dedicated to filtering high-quality multimodal news data. |
| Each input contains a news text and its associated news image. Evaluate the two modalities jointly. | |
| I. Judgment Criteria: | |
| 1. Applicable conditions (must be fully met to return “Usable”): | |
| (1) The text must possess a clear news format and complete structure. The selected content cannot be fragmented isolated words, but must have a clear description of the news event, and contain the basic elements constituting the news. |
| Dataset | Language | Original Quantity | Filtered Quantity |
|---|---|---|---|
| Fakeddit ( Nakamura et al., 2020 ) | English | 1,063,106 | 2,922 |
| FineFake ( Zhou et al., 2026 ) | English | 16,909 | 2,461 |
| MR2 ( Hu et al., 2023 ) | Chinese & English | 14,700 | 442 |
| Weibo-17 ( Jin et al., 2017 ) | Chinese | 9,528 | 942 |
| Weibo-21 ( Nan et al., 2021 ) | Chinese | 9,128 | 850 |
| CFND ( Zhang et al., 2024b ) | Chinese | 26,665 | 380 |
| Dataset | Samples | Candidate Pairs | Verified Event Groups | Grouped Samples |
|---|---|---|---|---|
| Weibo-21 | 4,493 | 286 | 74 | 198 |
| Fakeddit | 5,919 | 413 | 96 | 271 |
| SSS | 7,997 | 672 | 143 | 406 |
| Harm label | Level | Typical content |
|---|---|---|
| 0.1 | Very low harm | Entertainment gossip, minor misleading content, and everyday misinformation with a limited scope of influence. |
| 0.3 | Low harm | Content that may cause localized misunderstanding, but the impact is limited and the harm is controllable. |
| 0.5 | Moderate harm | Content involving public events, social controversies, or collective perceptions that may provoke disputes or cause a certain degree of adverse impact. |
| 0.7 | High harm | Content involving public safety, health, politics, disasters, or intergroup conflict that may produce a clear social impact. |
| 0.9 | Very high harm | Content that may induce panic, real-world harm, risks to public order, or large-scale collective impact. |
| Evaluation | Weibo-21 | Fakeddit | SSS |
|---|---|---|---|
| Human annotator agreement (Krippendorff’s ) | 0.9598 | 0.9303 | 0.9413 |
| LLM–human agreement (mean agreement) | 92.02% | 89.37% | 90.13% |
| Adjudication Trigger Rate | 0.78% (35/4,493) | 1.32% (78/5,919) | 1.25% (100/7,997) |
| Primary Category | Secondary Category | Retrieval Domains |
| Encyclopedias and Knowledge Bases | wikipedia, britannica, baidu | |
| Communities and Content Platforms | zhihu, douban, reddit, quora, youtube, weibo, sina, tencent, 163, sohu | |
| News Media | News Agencies | reuters, apnews, xinhuanet, chinanews, tass |
| Transnational Media | bbc, cnn, aljazeera, dw, euronews, france24, rt, sputniknews, chinadaily, sky, ifeng | |
| National Media | cctv, people, thepaper, huanqiu, guancha, gmw, nytimes, washingtonpost, theguardian, foxnews, cbsnews, nbcnews, npr, pbs, time, newsweek, independent, dailymail, usatoday, telegraph, express, mirror, huffpost, cbc, theglobeandmail, ndtv, thehindu, hindustantimes, indianexpress, indiatoday, straitstimes, japantimes, haaretz, timesofisrael, jpost, elpais, nzherald, buzzfeednews, vice, vox, slate, theatlantic, buzzfeed | |
| Local and Regional Media | bjnews, oeeee, scmp, chicagotribune, boston, bostonglobe, miamiherald, syracuse, denverpost, seattletimes, dallasnews, houstonchronicle, tampabay, kansascity, cleveland, oregonlive, sltrib, latimes, bangordailynews, nola, jacksonville, mercurynews, thestar, torontosun, vancouversun, smh, abc7, abc7ny, ktla, cbslocal, abc13, 9news, whyy | |
| Dataset | Temporal Subset | Candidates | After Explicit Filtering | After LLM Screening |
|---|---|---|---|---|
| Weibo-21 | Pre-news | 8,193 | 6,448 | 4,094 |
| Post-news | 9,894 | 7,322 | 4,247 | |
| Fakeddit | Pre-news | 11,727 | 8,456 | 6,757 |
| Post-news | 13,791 | 7,333 | 4,618 | |
| SSS | Pre-news | 18,526 | 12,324 | 8,606 |
| Post-news | 21,718 | 9,167 | 6,551 |
| System Prompt |
| You are a rigorous semantic stance annotator for news-evidence pairs. You need to judge the semantic relation between the evidence and the core event, subject, and claim expressed by the news. Please assign a continuous stance score in the interval to each evidence item. Interpret the score using the following ranges: • Strong refutation: the evidence directly conflicts with the core claim of the news, or explicitly provides facts opposite to the core content of the news. • Weak refutation: the evidence is partially inconsistent with the news, or weakly negates or weakens a key detail of the news. • Neutral: the evidence is only background-related or topic-related, provides insufficient information, cannot determine support or refutation, or is basically irrelevant to the news. • Weak support: the evidence is partially consistent with the news, or provides weak support for a key detail of the news. • Strong support: the evidence explicitly supports the core claim of the news, or provides facts highly consistent with the core content of the news. Important constraints: • Support or refutation only indicates the semantic relation between the evidence text and the news text, and is independent of the actual truthfulness of the news. • Do not complete missing facts using commonsense or external knowledge. Make the judgment only based on the input text. Output requirement: Return only one decimal number in , without explanations or additional text. |
| Score Range | Mapped Value | Meaning |
|---|---|---|
| 0.1 | Severe conflict in core facts or direct denial. | |
| 0.3 | Certain contradictions or inconsistencies exist, with weak refutation strength. | |
| 0.5 | Content is irrelevant or lacks sufficient information to support a judgment. | |
| 0.7 | Possesses partial factual consistency, providing limited support. | |
| 0.9 | Highly consistent at the key factual level, explicitly supporting the original claim. |
| Type | Parameter | Value |
| Feature | Text dim. | 768 |
| Image dim. | 768 | |
| Evidence dim. | 768 | |
| Projection dim. | 768 | |
| Relation dim. | 32 | |
| Graph | 0.6 |
| Role | Content |
|---|---|
| System | You are a rigorous multimodal news authenticity evaluator. Your task is to determine whether a news item is fake or real by jointly considering the news text, the associated image, and the retrieved external evidence. Decision criteria: (1) Judge whether the core claim expressed by the news is supported or contradicted by the visual content and the retrieved evidence. (2) Use the retrieved evidence only as contextual evidence. Do not assume that an evidence source is always correct, and do not use prior knowledge beyond the supplied inputs. (3) If the image-text pair or the evidence reveals clear factual contradiction, temporal inconsistency, entity mismatch, or unsupported fabrication, classify the news as fake. (4) If the news text, image, and evidence are mutually consistent and the evidence supports the core claim, classify the news as real. (5) If the evidence is insufficient, make the most likely binary judgment based on the supplied news text, image, and evidence. Output requirement: Return exactly one label: Fake or Real . Do not provide explanations. |
| User | News text: [NEWS TEXT] News image: [NEWS IMAGE] Retrieved evidence snippets: [EVIDENCE SNIPPETS] |
| MLLM | Fake or Real |
| Category | Method | Accuracy | Fake News | Real News | ||||
|---|---|---|---|---|---|---|---|---|
| Precision | Recall | F1 | Precision | Recall | F1 | |||
| w/o EK | CAFE | 0.815 | 0.850 | 0.774 | 0.810 | 0.785 | 0.858 | 0.820 |
| MRML | 0.903 | 0.881 | 0.935 | 0.908 | 0.928 | 0.869 | 0.898 | |
| Event-Radar | 0.881 | 0.873 | 0.896 | 0.884 | 0.889 | 0.865 | 0.877 | |
| MSACA | 0.894 | 0.861 | 0.944 | 0.900 | 0.935 | 0.841 | 0.886 | |
| w/ EK | KEHGNN-FD | 0.764 | 0.752 | 0.801 | 0.776 | 0.779 | 0.726 | 0.751 |
| Category | Method | Accuracy | Fake News | Real News | ||||
|---|---|---|---|---|---|---|---|---|
| Precision | Recall | F1 | Precision | Recall | F1 | |||
| w/o EK | CAFE | 0.786 | 0.773 | 0.864 | 0.816 | 0.807 | 0.691 | 0.745 |
| MRML | 0.860 | 0.852 | 0.900 | 0.876 | 0.870 | 0.811 | 0.839 | |
| Event-Radar | 0.840 | 0.831 | 0.888 | 0.859 | 0.852 | 0.780 | 0.815 | |
| MSACA | 0.846 | 0.839 | 0.889 | 0.864 | 0.855 | 0.794 | 0.823 | |
| w/ EK | KEHGNN-FD | 0.763 | 0.752 | 0.846 | 0.796 | 0.780 | 0.662 | 0.716 |
| Category | Method | Accuracy | Fake News | Real News | ||||
|---|---|---|---|---|---|---|---|---|
| Precision | Recall | F1 | Precision | Recall | F1 | |||
| w/o EK | CAFE | 0.726 | 0.709 | 0.802 | 0.753 | 0.750 | 0.644 | 0.693 |
| MRML | 0.758 | 0.742 | 0.821 | 0.779 | 0.780 | 0.690 | 0.732 | |
| Event-Radar | 0.754 | 0.755 | 0.779 | 0.767 | 0.752 | 0.726 | 0.739 | |
| MSACA | 0.762 | 0.757 | 0.801 | 0.778 | 0.769 | 0.721 | 0.744 | |
| w/ EK | KEHGNN-FD | 0.674 | 0.672 | 0.726 | 0.698 | 0.675 | 0.616 | 0.644 |
| Variation | Accuracy | Macro F1 | Fake F1 | HHF F1 | HHF Recall |
|---|---|---|---|---|---|
| w/o Evidence Graph | 0.913 | 0.913 | 0.915 | 0.956 | 0.961 |
| – w/o Edge Weights | 0.920 | 0.920 | 0.921 | 0.959 | 0.956 |
| – w/o Edge Types | 0.914 | 0.914 | 0.916 | 0.960 | 0.967 |
| w/o Harm-aware Branch | 0.915 | 0.915 | 0.914 | 0.946 | 0.919 |
| – w/o | 0.923 | 0.923 | 0.923 | 0.954 | 0.936 |
| – w/o | 0.924 | 0.924 | 0.924 | 0.947 | 0.932 |
| Variation | Accuracy | Macro F1 | Fake F1 | HHF F1 | HHF Recall |
|---|---|---|---|---|---|
| w/o Evidence Graph | 0.909 | 0.907 | 0.921 | 0.939 | 0.952 |
| – w/o Edge Weights | 0.915 | 0.914 | 0.924 | 0.936 | 0.935 |
| – w/o Edge Types | 0.911 | 0.910 | 0.921 | 0.930 | 0.935 |
| w/o Harm-aware Branch | 0.912 | 0.911 | 0.920 | 0.925 | 0.907 |
| – w/o | 0.909 | 0.908 | 0.917 | 0.926 | 0.919 |
| – w/o | 0.906 | 0.905 | 0.913 | 0.922 | 0.900 |
| Variation | Accuracy | Macro F1 | Fake F1 | HHF F1 | HHF Recall |
|---|---|---|---|---|---|
| w/o Evidence Graph | 0.822 | 0.818 | 0.844 | 0.892 | 0.956 |
| – w/o Edge Weights | 0.837 | 0.835 | 0.851 | 0.886 | 0.923 |
| – w/o Edge Types | 0.826 | 0.824 | 0.843 | 0.877 | 0.928 |
| w/o Harm-aware Branch | 0.839 | 0.838 | 0.850 | 0.878 | 0.905 |
| – w/o | 0.839 | 0.838 | 0.852 | 0.888 | 0.923 |
| – w/o | 0.833 | 0.832 | 0.844 | 0.892 | 0.917 |
| Category | Method | 0.1 | 0.3 | 0.5 | 0.7 | 0.9 | Overall |
|---|---|---|---|---|---|---|---|
| w/o EK | CAFE | 0.825 | 0.806 | 0.816 | 0.805 | 0.818 | 0.815 |
| MRML | 0.898 | 0.909 | 0.902 | 0.908 | 0.900 | 0.903 | |
| Event-Radar | 0.886 | 0.871 | 0.882 | 0.879 | 0.884 | 0.881 | |
| MSACA | 0.889 | 0.898 | 0.892 | 0.900 | 0.890 | 0.894 | |
| w/ EK | KEHGNN-FD | 0.771 | 0.756 | 0.769 | 0.755 | 0.765 | 0.764 |
| NSLM | 0.882 | 0.871 | 0.879 | 0.876 | 0.881 | 0.878 |
| Category | Method | 0.1 | 0.3 | 0.5 | 0.7 | 0.9 | Overall |
|---|---|---|---|---|---|---|---|
| w/o EK | CAFE | 0.792 | 0.782 | 0.789 | 0.783 | 0.786 | 0.786 |
| MRML | 0.852 | 0.865 | 0.858 | 0.863 | 0.860 | 0.860 | |
| Event-Radar | 0.844 | 0.836 | 0.842 | 0.837 | 0.839 | 0.840 | |
| MSACA | 0.840 | 0.849 | 0.844 | 0.850 | 0.847 | 0.846 | |
| w/ EK | KEHGNN-FD | 0.765 | 0.758 | 0.766 | 0.761 | 0.763 | 0.763 |
| NSLM | 0.844 | 0.855 | 0.847 | 0.854 | 0.849 | 0.850 |
| Category | Method | 0.1 | 0.3 | 0.5 | 0.7 | 0.9 | Overall |
|---|---|---|---|---|---|---|---|
| w/o EK | CAFE | 0.731 | 0.722 | 0.729 | 0.724 | 0.722 | 0.726 |
| MRML | 0.752 | 0.762 | 0.756 | 0.761 | 0.756 | 0.758 | |
| Event-Radar | 0.755 | 0.748 | 0.760 | 0.751 | 0.752 | 0.754 | |
| MSACA | 0.757 | 0.768 | 0.760 | 0.765 | 0.759 | 0.762 | |
| w/ EK | KEHGNN-FD | 0.680 | 0.668 | 0.677 | 0.671 | 0.674 | 0.674 |
| NSLM | 0.775 | 0.785 | 0.778 | 0.784 | 0.780 | 0.781 |
| Dataset | Predictor | MAE | RMSE |
|---|---|---|---|
| Weibo-21 | Text-only MLP | 0.103 | 0.137 |
| RAEGNet harm branch | 0.076 | 0.103 | |
| Fakeddit | Text-only MLP | 0.112 | 0.145 |
| RAEGNet harm branch | 0.084 | 0.112 | |
| SSS | Text-only MLP | 0.119 | 0.153 |
| RAEGNet harm branch | 0.091 | 0.119 |
| Dataset | Evaluation Set | Model | Accuracy | Macro F1 | Fake F1 | HHF F1 |
|---|---|---|---|---|---|---|
| Weibo-21 | Original | w/o Harm-aware Risk | 0.915 | 0.915 | 0.914 | 0.946 |
| Full Model | 0.931 | 0.931 | 0.934 | 0.972 | ||
| Harm-matched | w/o Harm-aware Risk | 0.902 | 0.902 | 0.903 | 0.928 | |
| Full Model | 0.914 | 0.914 | 0.916 | 0.946 | ||
| Fakeddit | Original | w/o Harm-aware Risk | 0.912 | 0.911 | 0.920 | 0.925 |
| Full Model | 0.919 | 0.918 | 0.926 | 0.947 |
| Dataset | Retrieval Query | Backbone | Accuracy | Macro F1 | Fake F1 | HHF F1 |
|---|---|---|---|---|---|---|
| Weibo-21 | Entity-level | Evidence Fusion | 0.889 | 0.889 | 0.891 | 0.932 |
| RAEGNet | 0.914 | 0.914 | 0.916 | 0.955 | ||
| Event-level | Evidence Fusion | 0.907 | 0.907 | 0.910 | 0.946 | |
| RAEGNet | 0.928 | 0.928 | 0.930 | 0.968 | ||
| Fakeddit | Entity-level | Evidence Fusion | 0.872 | 0.870 | 0.884 | 0.912 |
| RAEGNet | 0.899 | 0.898 | 0.908 | 0.934 |
| Dataset | Method | Evidence Source | Accuracy | Fake F1 | Real F1 |
|---|---|---|---|---|---|
| Weibo-21 | KEHGNN-FD | Original | 0.764 | 0.776 | 0.751 |
| ELERF | 0.801 | 0.812 | 0.789 | ||
| NSLM | Original | 0.878 | 0.885 | 0.870 | |
| ELERF | 0.896 | 0.901 | 0.890 | ||
| ERIC-FND | Original | 0.844 | 0.857 | 0.828 | |
| ELERF | 0.879 | 0.886 | 0.871 |