IVT-Guard: All-in-One Reasoning Model for AI-Generated Content Detection
Authors: Hongwei Niu, Yunpeng Luo, Hanjun Li, Ziyin Zhou, Jianghang Lin, Ke Yan, Shouhong Ding, Shengchuan Zhang, +1 more
Organizations: Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, Xiamen 361005, P.R. China · Tencent YouTu Lab
The rapid proliferation of highly realistic AI-Generated Content (AIGC) necessitates robust and interpretable detection mechanisms. However, existing detectors are predominantly confined to single modalities and provide binary outputs without reasoning. While Multimodal Large Language Models (MLLMs) present a promising solution, their development is constrained by the scarcity of multimodal reasoning data and the reasoning-detection optimization dilemma, where explicit reasoning supervision can compromise detection accuracy. To this end, we introduce IVT-Set, a comprehensive dataset comprising over 152K diverse image, video, and text samples equipped with multi-granularity Chain-of-Thought (CoT) reasoning trajectories. Based on it, we propose IVT-Guard, a pioneering framework for unified and interpretable AIGC detection across image, video, and text modalities. Furthermore, to overcome the aforementioned optimization dilemma, we design a novel three-stage training paradigm: Artifact-Aware Pre-training, Artifact-to-Evidence Supervised Fine-Tuning via artifact-aware injection, and Evidence-Verdict Consistency Group Relative Policy Optimization. Extensive experiments demonstrate that IVT-Guard achieves state-of-the-art detection performance across in-domain, out-of-domain, and cross-dataset settings while delivering faithful reasoning. Code and data will be released.
Figures & tables
Figure 1
Figure 2: Comparison of OOD accuracy (%) under different SFT strategies for three MLLMs on image (I), video (V), and text (T). Colors indicate accuracy changes relative to the standard CoT.
Figure 3: Data annotation pipeline for IVT-Set, consisting of CoT generation, multi-judge quality judgment, and iterative critique–refinement to produce high-quality multi-granularity CoT.
Figure 4: Overview of the three-stage training pipeline for IVT-Guard. Stage 1: AAP enhances sensitivity to modality-specific generation artifacts; Stage 2: A2E-SFT integrates artifact-aware representations into the MLLM for evidence-grounded reasoning; and Stage 3: EVC-GRPO jointly optimizes detection accuracy and evidence–verdict consistency.
Method
ID
CD
OOD
Avg.
Chameleon
LOKI
GenImage
GenBuster++
Real
FLUX.1 Krea
Qwen-Image
Midjourney v6
Infinity
Janus-Pro-7B
LlamaGen
Image-specific Detectors
UnivFD ( Ojha et al., 2023 )
77.75
52.42
67.70
72.97
54.70
82.75
79.88
68.00
73.38
73.62
74.12
87.38
72.24
DIRE ( Wang et al., 2023 )
89.00
70.92
71.23
75.95
72.35
89.75
72.88
55.75
74.75
81.50
71.00
87.25
79.25
NPR ( Tan et al., 2024 )
84.62
60.15
66.07
75.70
52.90
91.25
61.12
54.62
70.12
81.00
62.62
65.25
72.58
AIDE ( Yan et al., 2025 )
88.62
61.81
77.74
87.76
60.15
89.25
79.75
74.12
74.25
84.12
80.62
90.00
80.74
Table 2: Accuracy comparison (%) on the image modality of IVT-Set and four cross-dataset benchmarks. Avg. denotes the mean of ID accuracy, average CD accuracy, and average OOD accuracy. Best and second-best results are shown in bold and underlined , respectively.
Method
ID
CD
OOD
Avg.
LOKI
AIGVDBench †
GenBuster++
Real
EasyAnimateV5.1
HunyuanVideo-I2V
Kling 2.1
Veo 3
Vidu Q2
Video-specific Detectors
ReStraV ( Internò et al., 2025 )
69.25
36.64
44.79
52.00
77.25
48.62
67.62
58.63
50.62
54.37
57.75
NSG-VD ( Zhang et al., 2025 )
50.38
57.46
78.73
50.40
18.75
46.00
45.25
50.12
48.00
47.88
51.75
DeMamba ( Chen et al., 2026b )
87.62
79.25
44.22
62.45
88.25
67.00
76.00
63.88
55.88
74.38
73.50
Generic MLLMs
Table 3: Accuracy comparison (%) on the video modality of IVT-Set and three cross-dataset benchmarks. † denotes samples from four closed-source generators in AIGVDBench ( Ma et al., 2026 ) . Avg. denotes the mean of ID accuracy, average CD accuracy, and average OOD accuracy. Best and second-best results are shown in bold and underlined , respectively.
Figure 5: Ablation study of visual encoder combinations in Artifact-Aware Pre-training.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Method
ID
CD
OOD
Avg.
LOKI
BiScope
DetectRL-X
AI Detector Bench
Real
Claude Opus 4.1
Gemini 2.5 Pro
GPT-4o
Text-specific Detectors
ModernBERT ( Warner et al., 2025 ; Drayson et al., 2025 )
51.75
78.16
63.00
54.62
62.10
89.00
64.38
62.88
66.25
62.28
Desklib AI ( Desklib, 2025 )
79.75
93.75
96.61
84.28
83.40
91.00
93.50
91.62
91.88
87.09
Generic MLLMs & LLMs
Qwen3-VL-8B
43.25
51.81
42.14
50.93
50.70
69.50
47.62
45.38
44.25
47.94
Appendix
Table 6: Performance comparison (Acc, %) on the text modality of IVT-Set and three cross-dataset benchmarks. Results are reported under in-domain (ID), cross-dataset (CD), and out-of-domain (OOD) settings. Avg. denotes the average of the ID accuracy, the mean CD accuracy, and the mean OOD accuracy. Best results are in bold and the second-best are underlined .
Method
Image
Video
Text
Score
Elo
Score
Elo
Score
Elo
GLM-4.6V
3.09
1036
2.91
964
3.33
1132
Qwen3-VL-235B-A22B
3.88
1352
3.06
1024
3.68
1294
Qwen3-VL-8B-Instruct
3.91
1364
2.96
984
3.41
1164
IVT-Guard (Ours)
4.70
1676
4.34
1564
4.28
1712
Appendix
Table 7: Explanation reliability evaluation on the OOD test set. We report the absolute score evaluated by Qwen3.5-122B-A10B and the pairwise Elo rating obtained from human evaluation for each modality, with all methods initialized at an Elo rating of 1,000.
Training Data
ID
OOD
Image (Acc/F1)
Video (Acc/F1)
Text (Acc/F1)
Image (Acc/F1)
Video (Acc/F1)
Text (Acc/F1)
Image only
98.00/98.31
51.38/42.62
47.00/36.45
91.13/91.03
48.45/39.05
48.21/36.95
Video only
68.00/67.54
91.25/91.25
45.38/39.43
66.73/63.79
75.73/75.14
46.54/39.80
Text only
56.25/51.55
49.88/46.38
88.63/88.62
55.56/50.97
49.05/46.77
88.33/88.30
Image + Video
98.75/98.75
92.63 / 92.62
46.00/39.73
91.58/91.52
77.23/76.63
46.88/39.25
Video + Text
71.63/71.11
91.75/91.75
97.75/97.75
70.90/68.34
78.65 / 78.16
96.04/96.12
Appendix
Table 8: Cross-modal evaluation and modality ablation at the A2E-SFT stage. Results are reported as ID/OOD accuracy and F1 (%). Best results are in bold .
Method
Perturbation
ID
OOD
Image (Acc/F1)
Video (Acc/F1)
Text (Acc/F1)
Image (Acc/F1)
Video (Acc/F1)
Text (Acc/F1)
All
None
98.75/98.75
92.00/91.99
98.38/98.37
93.48/93.43
82.80/82.71
98.00/98.00
Image
JPEG Compression (QF=90)
98.38/98.37
92.00/91.99
98.38/98.37
93.69/93.64
82.80/82.71
98.00/98.00
Gaussian Blur ( σ=1.0 )
96.63/96.62
92.00/91.99
98.38/98.37
92.69/92.65
82.80/82.71
98.00/98.00
Video
Shuffle frame
98.75/98.75
92.25/92.24
98.38/98.37
93.48/93.43
82.55/82.45
98.00/98.00
Resize ×0.7
98.75/98.75
89.13/89.06
98.38/98.37
93.48/93.43
81.93/81.87
98.00/98.00
Appendix
Table 9: Robustness evaluation of IVT-Guard under modality-specific input perturbations. Results are reported as accuracy/F1 (%) under in-domain (ID) and out-of-domain (OOD) settings for image, video, and text modalities.
Figure 6: OOD accuracy across image, video, and text modalities for three model families at different training stages. Bars denote OOD accuracy (%), while lines denote mean recall (%) across OOD generators, with “fake” treated as the positive category.
Figure 7: Overview of the IVT-Set construction pipeline, including modality-specific data pairing and filtering, diverse source and category coverage, and multi-granularity CoT annotations.
Method
Model Type
Resolution
Data Scale
Data Split
BigGAN ( Brock et al., 2019 )
GAN
200 × 200 / 128 × 128
742
ID
StyleGAN ( Karras et al., 2019 )
GAN
200 × 200
1,155
ID
DF-GAN ( Tao et al., 2022 )
GAN
256 × 256
198
ID
GALIP ( Tao et al., 2023 )
GAN
224 × 224
994
ID
GigaGAN ( Kang et al., 2023 )
GAN
512 × 512
497
ID
Stable Diffusion v1.4 ( CompVis, 2022 )
Diffusion
1024 × 704 / 1024 × 768 704 × 1024 / 768 × 1024
767
ID
Appendix
Table 10: Details of the image modality generation methods in IVT-Set. The dataset includes 20 generators spanning GAN, diffusion/flow-based, and autoregressive paradigms, including a closed-source system whose architecture is not publicly disclosed. Generators are split into in-domain (ID) and out-of-domain (OOD) to evaluate cross-generator generalization. Image sizes are reported as width × height in pixels. For generators with mixed-resolution samples, we report up to the four most common image sizes in descending order.
Method
Duration (s)
Resolution
FPS
Data Scale
Data Split
Pika ( labs, 2022 )
3.00–4.00
768p / 576p 640p / 1024p
12–24
1,000
–
SVD ( Blattmann et al., 2023 )
3.57
576p
7
1,000
–
DynamiCrafter ( Xing et al., 2024 )
2.00
576p
8
1,000
–
OpenSora ( Zheng et al., 2024 )
2.00
256p / 512p
8
1,000
–
CogVideoX ( Yang et al., 2025b )
5.04
1024p
24
1,000
–
LTX-Video ( HaCohen et al., 2024 )
1.37–10.25
672p / 480p 512p / 240p
6–60
4,999
ID
Appendix
Table 11: Details of the video modality generation methods in IVT-Set. The dataset includes 14 generators with diverse resolutions and frame rates. Generators are split into in-domain (ID) and out-of-domain (OOD) categories when applicable to evaluate cross-generator generalization. Video resolutions are reported as output heights in pixels. For generators with mixed-resolution samples, we report up to the four most common output heights in descending order.
Method
R./F. Ratio
Length
Data Scale
Data Split
Llama-3.3-70B-Instruct ( Grattafiori et al., 2024 )
1.09
20–1,190
1,486
ID
gpt-oss-120b ( Agarwal et al., 2025 )
0.99
27–1,056
2,635
ID
GLM-4.5-Air ( Zeng et al., 2025 )
1.04
20–1,116
2,469
ID
DeepSeek-V3.1 ( Liu et al., 2024a )
1.09
20–1,516
2,473
ID
Kimi K2-Instruct-0905 ( Team et al., 2025 )
1.05
20–787
2,190
ID
Qwen3-235B-A22B-Instruct-2507 ( Yang et al., 2025a )
0.99
21–1,206
2,680
ID
Appendix
Table 12: Details of the text modality generation methods in IVT-Set. The dataset includes 10 representative LLMs covering both open-source and closed-source models. Models are split into in-domain (ID) and out-of-domain (OOD) categories to evaluate cross-model generalization. R./F. Ratio denotes the total real-text length divided by the total corresponding generated-text length. Length reports the minimum–maximum generated-text length in words.
Figure 8: Data statistics of the image modality in IVT-Set. (a) Category distribution of all images across 13 semantic categories. (b) Generator distribution of fake images across 20 models spanning GAN, Diffusion, and Autoregressive paradigms. (c) Comparison of the mean quality scores of fake and real images in each category, showing that the quality scores of the two are very close across all categories.
Figure 9: Data statistics of the video modality in IVT-Set. (a) Category distribution of all videos across 11 semantic categories. (b) Generator distribution of fake videos across 14 generation models, covering both open-source and commercial models. (c) Correlation between aesthetic score and temporal consistency score for fake and real videos, showing that the two distributions largely overlap.
Figure 10: Data statistics of the text modality in IVT-Set. (a) Category distribution across 7 domains. (b) Generator distribution of fake texts from 10 LLMs. (c) Stylometric similarity among real-text sources. (d) Stylometric similarity among fake-text generators. For (c) and (d), each source is represented by a category-averaged profile of 20 rule-based stylometric features. Pairwise similarity is computed as the Pearson correlation between standardized source profiles and rescaled to [0,1] , with higher values indicating more similar profiles.
Figure 11: Prompt template for CoT Generation on real images .
Figure 12: Prompt template for CoT Generation on fake images .
Figure 13: Prompt template for CoT Generation on real videos.
Figure 14: Prompt template for CoT Generation on fake videos .
Figure 15: Prompt template for CoT Generation on real texts .
Figure 16: Prompt template for CoT Generation on fake texts .
Figure 17: Prompt template for CoT Quality Judgment.
Figure 18: Prompt template for Critique Refinement on real images .
Figure 19: Prompt template for Critique Refinement on fake images .
Figure 20: Prompt template for Critique Refinement on real videos .
Figure 21: Prompt template for Critique Refinement on fake videos .
Figure 22: Prompt template for Critique Refinement on real texts .
Figure 23: Prompt template for Critique Refinement on fake texts .
The rapid proliferation of AI-Generated Images (AIGIs) poses severe misinformation risks, making AIGI detection critical yet challenging. Traditional detection paradigms mainly rely on low-level features, whereas recent research increasingly focuses on leveraging the general understanding ability of Multimodal Large Language Models (MLLMs) to achieve better generalization, yet it still suffers from limited extensibility and expensive data annotations. Instead of building yet another detector, we recast AIGI detection as learned, reasoning-based evidence synthesis over a pool of heterogeneous off-the-shelf detectors, realized through EvoGuard, a novel agentic framework. A capability-aware selection mechanism profiles each detector and gathers complementary evidence per sample; a dynamic orchestration mechanism then reasons over heterogeneous outputs across multiple rounds, cross-validating conflicting or low-confidence signals before concluding. This design exploits the complementary strengths among heterogeneous detectors, transcending the limits of any single model. Furthermore, optimized by a GRPO-based Agentic Reinforcement Learning algorithm using only low-cost binary labels, it eliminates the reliance on fine-grained annotations. Extensive experiments demonstrate that this learned reasoning paradigm outperforms single-detector and static ensembling, achieving SOTA accuracy while mitigating the bias between positive and negative samples. More importantly, it allows the plug-and-play integration of new detectors to boost overall performance in a train-free manner, offering a highly practical, long-term solution to ever-evolving AIGI threats. Source code will be publicly available upon acceptance.
Chenyang Zhu, Maorong Wang, Jun Liu +2
The University of Tokyo Tokyo, Japan · National Institute of Informatics Tokyo, Japan
The rapid advancement and widespread adoption of Large Language Models (LLMs) have elevated the need for reliable AI-generated content (AIGC) detection, which remains challenging as models evolve. We introduce AIGC-text-bank, a comprehensive multi-domain dataset with diverse LLM sources and authorship scenarios, and propose REVEAL, a detection framework that generates interpretable reasoning chains before classification. Our approach uses a two-stage training strategy: supervised fine-tuning to establish reasoning capabilities, followed by reinforcement learning to improve accuracy, improve logical consistency, and reduce hallucinations. Extensive experiments show that REVEAL achieves state-of-the-art performance across multiple benchmarks, offering a robust and transparent solution for AIGC detection. The project is open-source at https://aka.ms/reveal
Zhao Wang, Max Xiong, Jianxun Lian +1
Gaoling School of Artificial Intelligence, Renmin University of China · Duke University · Microsoft Research Asia
AI-generated content (AIGC) is rapidly improving, creating an urgent need for detectors that generalize across data sources, deployment pipelines, and visual modalities. A strongly generalizable detector should remain robust under distributional variations. However, we identify a consistent failure mode: SOTA AI-generated image detectors often collapse when applied to frames extracted from videos. Through systematic analysis, we show that this cross-modal gap arises from both entangled synthesis-agnostic video processing shifts, including color conversion, codec compression, resizing, and blur, and model-specific fingerprints introduced by modern video generators. Motivated by these findings, we propose VINA (Video as Natural Augmentation), a unified AIGC detection framework that jointly trains on image and video data. VINA uses video frames as physically grounded natural augmentations and further introduces a cross-modal supervised contrastive objective to align image and video representations under a shared real/fake decision boundary. Extensive experiments on 14 image, video, and in-the-wild benchmarks show that VINA delivers bidirectional gains, improves robustness and transferability, and achieves state-of-the-art performance across nearly all evaluated settings without complex augmentation or dataset-specific tuning.
Zhengcen Li, Chenyang Jiang, Liangxu Su +4
Harbin Institute of Technology, Shenzhen · Shenzhen Loop Area Institute · Pengcheng Laboratory