REVEAL: Robust Evolution of Vision-Language Models for Explainable AI-Video Detection
Organizations: Columbia University · Rutgers University
Abstract
The rapid advancement of AI-generated video poses challenges to digital authenticity and security. Current detection methods, often trained on specific datasets, struggle with the ever-evolving landscape of generative techniques and unseen manipulations. We introduce a framework leveraging Vision Language Models (VLMs) for robust AI-generated video detection. Our approach equips the VLM with the ability to reason about video content and use external tools to identify subtle inconsistencies, mirroring human system 2 thinking. Our self-evolving VLM dynamically selects and composes appropriate tools, enhancing its ability to generalize to novel video generation techniques. The modular design promotes interpretability, allowing for a clearer understanding of VLM's decision-making process. To evaluate, we establish the first benchmark VidForensic containing 1.4k+ high-quality AI-generated videos across eight generative models. Experiments show that REVEAL improves F1 scores by 9.1% to 30.2% over top baselines across our datasets for VLMs, notably for GPT-4o, Gemini 1.5 pro, and QWen-VL-Max, and Llava-One-Vision-7B. While open-world AI-video detection remains an open challenge, our results indicate that existing methods fail primarily because they lack tool-enabled, higher-order reasoning.
Figures & tables
| Dataset Source | Video Source | Type | # Videos | Res. | FPS | Length |
| PANDA-70M [ 15 ] | Youtube | Real | 200 | - | - | 1 10s |
| VidProM [ 45 ] | Text2Video-Zero [ 24 ] | AI | 200 | 512*512 | 4 | 2s |
| VideoCrafter2 [ 12 ] | AI | 200 | 512*320 | 10 | 1s | |
| ModelScope [ 42 ] | AI | 200 | 256*256 | 8 | 2s | |
| Pika [ 36 ] | AI | 200 | - | 24 | 3s | |
| Self-Collected | Youtube | Real | 45 | - | 30 | 1 4s |
| Category | Explicit Knowledge (EK) Toolkits |
| Appearance | Saturation, Denoised, Sharpen, Enhance, Segmentation Map |
| Motion | Optical flow, Landmark |
| Geometry | Depth map, Edge |
| LVLM | Method | VidForensic (VidProM) [ 45 ] | VidForensic (Self-collected) | Avg. | ||||||
| Pika [ 36 ] | T2vz [ 24 ] | Vc2 [ 12 ] | Ms [ 42 ] | OpenSORA [ 50 ] | Gen3 [ 40 ] | Kling [ 25 ] | SORA [ 6 ] | |||
| — | XGBoost [ 14 ] | 56.00/46.55 | 50.50/34.42 | 57.00/48.51 | 52.00/37.98 | 53.00/40.25 | 54.00/42.43 | 55.00/44.53 | 51.11/29.37 | 53.58/40.50 |
| — | SVM [ 16 ] | 63.00/56.65 | 78.00/77.68 | 86.00/85.90 | 85.50/85.43 | 67.50/63.94 | 65.00/60.02 | 68.50/65.43 | 52.22/26.48 | 70.72/65.19 |
| — | ViT [ 41 ] | 76.56/76.01 | 88.90/88.51 | 91.52/90.73 | 91.52/90.73 | 81.33/81.29 | 83.34/83.33 | 75.51/74.76 | 61.14/61.08 | 81.23/80.81 |
| Llava-OV-7B [ 28 ] | Direct prompt 1 | 53.50/14.68 | 61.00/37.10 | 61.00/37.10 | 58.50/30.25 | 52.50/12.11 | 50.00/1.96 | 50.00 /1.96 | 54.44/16.33 | 55.12/18.94 |
| Direct prompt 2 | 50.50/1.98 | 51.00/3.92 | 51.50/5.83 | 53.50/13.08 | 52.00/7.69 | 50.00/0.00 | 50.00/0.00 | 50.00/0.00 | 51.06/4.06 | |
| Model | Land. | Depth | Enhan. | Edge | Sharp. | Denoise | OPflow | Sat. | SAM |
| Llava-OV-7B [ 28 ] | ✓ | ✓ | ✓ | ||||||
| Qwen-VL-Max [ 37 ] | ✓ | ✓ | ✓ | ✓ | |||||
| Gemini-1.5-pro [ 17 ] | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | |||
| GPT-4o [ 33 ] | ✓ | ✓ | ✓ |
| Method | Trainset | Celeb-DF-v1 | |
| Acc. | F1 | ||
| Guo et al. [ 19 ] | FF++ [ 39 ] | 73.19 | – |
| RECCE [ 7 ] | FF++ [ 39 ] | 71.81 | – |
| MAT [ 49 ] | FF++ [ 39 ] | 71.81 | – |
| Baseline (Gemini-1.5-pro) | – | 44.00 | 17.65 |
| Baseline (GPT-4o) | – | 64.95 | 74.24 |
| Name | Input Text Token | 1080p Input Frame | Input Video |
| GPT4o-0806 | $2.50/1M | $2.763/1K | $22.215/1K |
| Gemini-1.5-pro-002 | $1.25/1M | $0.323/1K | $2.639/1K |
| Qwen-VL-MAX-0809 | $2.78/1M | $3.451/1K | $27.728/1K |
| Category | EK Name | EK Description (Summarized from LVLM) |
| Appearance | Saturation | AI-generated videos may exhibit anomalies in color rendering. Saturation estimation detects color unevenness, oversaturation, or undersaturation to identify artificial elements. |
| Denoised | Denoising isolates unnatural noise patterns present in AI-generated videos. Residual artifacts after denoising can signal synthesized or forged content. | |
| Sharpen | Sharpening frames emphasizes edges, making it easier to spot unnatural boundaries or blending artifacts, which may indicate forgery. | |
| Enhance | Image enhancement boosts details and contrast, revealing synthetic artifacts like unnatural textures or color inconsistencies. | |
| Segmentation Map | Segmentation maps identify mismatched regions in synthesized content, such as areas where the object segmentation boundaries do not align with real-world logic. | |
| Motion | Optical Flow | AI-generated videos may have abnormal motion patterns, such as discontinuous movements or unnatural trajectories. Optical flow estimation detects whether object motion in the video is smooth and adheres to physical laws. |
| LVLM | Method | VidForensic (VidProM) [ 45 ] | VidForensic (Self-collected) | Avg. | ||||||
| Pika [ 36 ] | T2vz [ 24 ] | Vc2 [ 12 ] | Ms [ 42 ] | OpenSORA [ 50 ] | Gen3 [ 40 ] | Kling [ 25 ] | SORA [ 6 ] | |||
| Qwen-VL-Max [ 37 ] | Baseline1 (w/o SP) | 72.50/63.09 | 75.00/67.53 | 82.00/78.57 | 76.00/69.23 | 67.50/53.24 | 62.00/40.62 | 54.50/19.47 | 58.89/39.34 | 68.55/51.24 |
| Baseline2 (w/o SP) | 60.50/38.76 | 75.00/68.35 | 71.50/62.25 | 72.50/64.05 | 60.50/38.76 | 52.00/14.29 | 50.00/7.41 | 56.67/26.42 | 62.33/39.56 | |
| Baseline3 (w/o SP) | 74.00 / 67.90 | 79.00 /75.58 | 84.50 / 83.06 | 79.50 /76.30 | 69.50/60.13 | 65.50/52.41 | 54.00/24.59 | 61.11/47.76 | 70.89/60.97 | |
| REVEAL (w/o SP) | 87.00 / 88.39 | 81.50 / 82.63 | 86.00 / 87.39 | 77.00/ 77.45 | 79.00 / 79.81 | 82.50 / 83.72 | 60.00 / 52.94 | 67.78 / 71.84 | 77.60 / 76.08 | |
| w/ video-specific Sel. | 70.14/62.83 | 78.50/ 76.76 | 82.25/81.38 | 80.17 / 78.70 | 77.25 / 74.48 | 69.44 / 61.53 | 70.27 / 62.65 | 74.02 / 69.99 | 75.26 / 71.04 | |