The rapid advancement of AI-generated video poses challenges to digital authenticity and security. Current detection methods, often trained on specific datasets, struggle with the ever-evolving landscape of generative techniques and unseen manipulations. We introduce a framework leveraging Vision Language Models (VLMs) for robust AI-generated video detection. Our approach equips the VLM with the ability to reason about video content and use external tools to identify subtle inconsistencies, mirroring human system 2 thinking. Our self-evolving VLM dynamically selects and composes appropriate tools, enhancing its ability to generalize to novel video generation techniques. The modular design promotes interpretability, allowing for a clearer understanding of VLM's decision-making process. To evaluate, we establish the first benchmark VidForensic containing 1.4k+ high-quality AI-generated videos across eight generative models. Experiments show that REVEAL improves F1 scores by 9.1% to 30.2% over top baselines across our datasets for VLMs, notably for GPT-4o, Gemini 1.5 pro, and QWen-VL-Max, and Llava-One-Vision-7B. While open-world AI-video detection remains an open challenge, our results indicate that existing methods fail primarily because they lack tool-enabled, higher-order reasoning.
Figures & tables
Figure 1 : An example of an AI-generated video from Kling [ 25 ] where REVEAL makes a correct prediction with explicit knowledge selection and prompt self-adaptation. REVEAL facilitates VLMs for video detection by calling explicit knowledge tools to extract useful information from the original videos and by providing structure-formatted output.
Figure 2 : We show the main pipeline of REVEAL for video detection. First, VLMs suggest relevant video detection tools. Next, based on the weighted accuracy, F1 , and the self-assessment score, SSA , each tool provides, we assemble a customized toolkit for each video detection VLM. In Fig. 4 , we show details of the self-adaptation of the structured prompt. The prompt tuning will be based on the VLM itself. Component marked with the logo are developed with the VLM-like GPT-4o [ 33 ] .
Dataset Source
Video Source
Type
# Videos
Res.
FPS
Length
PANDA-70M [ 15 ]
Youtube
Real
200
-
-
1 ∼ 10s
VidProM [ 45 ]
Text2Video-Zero [ 24 ]
AI
200
512*512
4
2s
VideoCrafter2 [ 12 ]
AI
200
512*320
10
1s
ModelScope [ 42 ]
AI
200
256*256
8
2s
Pika [ 36 ]
AI
200
-
24
3s
Self-Collected
Youtube
Real
45
-
30
1 ∼ 4s
Table 1 : Composition of the VidForensic. We collect high-quality video from multiple sources. For dataset source own-generated, we generate text-to-video samples with generators conditioned on text prompts collected from PANDA-70M [ 15 ] by ourselves.
Table 2 : Categories of explicit knowledge toolkits. Though all tools are proposed by LVLMs, we list and categorize all explicit knowledge that we collect from LVLM in the process of initial toolkit preparation into three VR categories.
Figure 3 : Prompt example for LVLM
Figure 4 : The flow of self-adaptation prompting process in REVEAL . Component marked with the logo are developed with the LVLM like GPT-4o [ 33 ] .
LVLM
Method
VidForensic (VidProM) [ 45 ]
VidForensic (Self-collected)
Avg.
Pika [ 36 ]
T2vz [ 24 ]
Vc2 [ 12 ]
Ms [ 42 ]
OpenSORA [ 50 ]
Gen3 [ 40 ]
Kling [ 25 ]
SORA [ 6 ]
—
XGBoost [ 14 ]
56.00/46.55
50.50/34.42
57.00/48.51
52.00/37.98
53.00/40.25
54.00/42.43
55.00/44.53
51.11/29.37
53.58/40.50
—
SVM [ 16 ]
63.00/56.65
78.00/77.68
86.00/85.90
85.50/85.43
67.50/63.94
65.00/60.02
68.50/65.43
52.22/26.48
70.72/65.19
—
ViT [ 41 ]
76.56/76.01
88.90/88.51
91.52/90.73
91.52/90.73
81.33/81.29
83.34/83.33
75.51/74.76
61.14/61.08
81.23/80.81
Llava-OV-7B [ 28 ]
Direct prompt 1
53.50/14.68
61.00/37.10
61.00/37.10
58.50/30.25
52.50/12.11
50.00/1.96
50.00 /1.96
54.44/16.33
55.12/18.94
Direct prompt 2
50.50/1.98
51.00/3.92
51.50/5.83
53.50/13.08
52.00/7.69
50.00/0.00
50.00/0.00
50.00/0.00
51.06/4.06
Table 3 : Performance comparison of REVEAL with three supervised learning-based detection methods and three direct prompting methods across eight datasets. For each dataset except SORA, we combine the real dataset from Panda-70M & AI-generated dataset together. For SORA, we combine them with 45 YouTube videos that collected by ourselves. We use four representative LVLMs to serve as the detector in our framework, including Llava-OV-7B [ 28 ] , Qwen-VL-Max [ 37 ] , Gemini-1.5-pro [ 17 ] , and GPT-4o [ 33 ] . The results are presented as Accuracy / F1-score in each cell. Numbers in bold show the top-1 best results, and numbers with underlined show the top-2 best results.
Model
Land.
Depth
Enhan.
Edge
Sharp.
Denoise
OPflow
Sat.
SAM
Llava-OV-7B [ 28 ]
✓
✓
✓
Qwen-VL-Max [ 37 ]
✓
✓
✓
✓
Gemini-1.5-pro [ 17 ]
✓
✓
✓
✓
✓
✓
GPT-4o [ 33 ]
✓
✓
✓
Table 4 : Summarization of model-specific explicit knowledge tools. The self-evolving selection strategy in REVEAL helps different LVLMs to customize their own useful tools for video detection.
Figure 5 : Reasoning visualization for REVEAL . We visualize three AI-generated samples and one real-world sample. The baseline reasoning illustrates the output reasoning when we directly query the LVLM with raw video and direct prompt. Our reasoning represents the output reasoning when we query LVLM with multiple explicit knowledges and structured prompt.
Method
Trainset
Celeb-DF-v1
Acc.
F1
Guo et al. [ 19 ]
FF++ [ 39 ]
73.19
–
RECCE [ 7 ]
FF++ [ 39 ]
71.81
–
MAT [ 49 ]
FF++ [ 39 ]
71.81
–
Baseline (Gemini-1.5-pro)
–
44.00
17.65
Baseline (GPT-4o)
–
64.95
74.24
Table 5 : Performance comparison of existing Deepfake detection baselines, the baseline prompts, and REVEAL on Celeb-DF-v1. Video-level accuracy (Acc.) and F1-score (F1) are used as evaluation metrics where available. The reported performance of RECCE and MAT are referenced from [ 44 ] .
Name
Input Text Token
1080p Input Frame
Input Video
GPT4o-0806
$2.50/1M
$2.763/1K
$22.215/1K
Gemini-1.5-pro-002
$1.25/1M
$0.323/1K
$2.639/1K
Qwen-VL-MAX-0809
$2.78/1M
$3.451/1K
$27.728/1K
Table 6 : Computation cost of REVEAL
Figure 6 : Dataset collection pipeline for VidForensic. Component marked with the logo are developed with the LVLM like GPT-4o [ 33 ] .
Category
EK Name
EK Description (Summarized from LVLM)
Appearance
Saturation
AI-generated videos may exhibit anomalies in color rendering. Saturation estimation detects color unevenness, oversaturation, or undersaturation to identify artificial elements.
Denoised
Denoising isolates unnatural noise patterns present in AI-generated videos. Residual artifacts after denoising can signal synthesized or forged content.
Sharpen
Sharpening frames emphasizes edges, making it easier to spot unnatural boundaries or blending artifacts, which may indicate forgery.
Enhance
Image enhancement boosts details and contrast, revealing synthetic artifacts like unnatural textures or color inconsistencies.
Segmentation Map
Segmentation maps identify mismatched regions in synthesized content, such as areas where the object segmentation boundaries do not align with real-world logic.
Motion
Optical Flow
AI-generated videos may have abnormal motion patterns, such as discontinuous movements or unnatural trajectories. Optical flow estimation detects whether object motion in the video is smooth and adheres to physical laws.
Table 7 : Details for nine explicit knowledge tools
LVLM
Method
VidForensic (VidProM) [ 45 ]
VidForensic (Self-collected)
Avg.
Pika [ 36 ]
T2vz [ 24 ]
Vc2 [ 12 ]
Ms [ 42 ]
OpenSORA [ 50 ]
Gen3 [ 40 ]
Kling [ 25 ]
SORA [ 6 ]
Qwen-VL-Max [ 37 ]
Baseline1 (w/o SP)
72.50/63.09
75.00/67.53
82.00/78.57
76.00/69.23
67.50/53.24
62.00/40.62
54.50/19.47
58.89/39.34
68.55/51.24
Baseline2 (w/o SP)
60.50/38.76
75.00/68.35
71.50/62.25
72.50/64.05
60.50/38.76
52.00/14.29
50.00/7.41
56.67/26.42
62.33/39.56
Baseline3 (w/o SP)
74.00 / 67.90
79.00 /75.58
84.50 / 83.06
79.50 /76.30
69.50/60.13
65.50/52.41
54.00/24.59
61.11/47.76
70.89/60.97
REVEAL (w/o SP)
87.00 / 88.39
81.50 / 82.63
86.00 / 87.39
77.00/ 77.45
79.00 / 79.81
82.50 / 83.72
60.00 / 52.94
67.78 / 71.84
77.60 / 76.08
w/ video-specific Sel.
70.14/62.83
78.50/ 76.76
82.25/81.38
80.17 / 78.70
77.25 / 74.48
69.44 / 61.53
70.27 / 62.65
74.02 / 69.99
75.26 / 71.04
Table 8 : Performance comparison of baselines and REVEAL with and without video-specific tool selection on eight datasets. For each dataset except SORA, we mix the real dataset from Panda-70M & AI-generated dataset together. For SORA, we mix it with 45 youtube videos that collected by ourselves. We use three representative LVLMs, including Qwen-VL-Max [ 37 ] , Gemini-1.5-pro [ 17 ] , and GPT-4o [ 33 ] . The results are presented as Accuracy / F1-score in each cell. Numbers in bold show the top-1 best results, and numbers with underlined show the top-2 best results.
School of Cyber Science and Engineering, Southeast University, Nanjing 210096, China · Purple Mountain Laboratories, Nanjing 210000, China · Engineering Research Center of Blockchain Application, Supervision And Management (Southeast University), Ministry of Education, China +2