CCDF: A Benchmark Dataset for Deepfake Detection in Real-World Surveillance Footage
Organizations: da/sec – Biometrics and Security Research Group, Hochschule Darmstadt · Computer Science and Engineering Department, Michigan State University · STARS team, Inria Center at Université Côte d’Azur
Abstract
Due to rapid advances in Generative AI, commercial video generation tools can be used to produce fabricated surveillance footage that can fool both human viewers and automated synthetic video detectors. Since these tools are so widely accessible, a malicious user can create a harmful video clip at minimal cost. The production and dissemination of such videos in high-stakes settings, such as crime reporting and elections, can misdirect emergency response efforts or distort political discourse. Existing deepfake video datasets, used by the research community to develop deepfake detection algorithms, exhibit two limitations: (1) they emphasize benign web content rather than footage of possibly malicious activity, and (2) they rely on older or open-source generators that do not represent recent advances in generative systems. We assemble CCtv DeepFakes (CCDF), a video deepfake dataset, to address both gaps. CCDF contains 1840 videos (460 real and 1380 generated) spanning 16 crime and accident categories, with generated content produced using three leading commercial systems: Grok Imagine, Google VEO 3.1, and OpenAI Sora 2. CCDF is a highly realistic, small-scale, manually annotated dataset targeting evaluation of detection models. We release three versions of the dataset: the raw generated data, a cleaned version in which video metadata are standardized between real and synthetic samples to prevent detectors from exploiting trivial cues, and an altered version simulating low-effort post-processing attacks. We evaluate CCDF with ten recent state-of-the-art detectors covering different detection approaches. Our results suggest that these approaches do not reliably distinguish CCDF's generated videos from real ones, despite their strong reported performance on existing datasets. These results further confirm that existing datasets are not well-suited to evaluating certain threats.
Figures & tables
| Dataset | T | # Samples | Generators | Content Style |
|---|---|---|---|---|
| DFDC Dolhansky et al. (2020) | V | 128K | 8 Open-Source Models | Face |
| DF40 Yan et al. (2024) | I/V | 1M | 40 Open-Source Models | Face |
| WildDeepfake Zi et al. (2020) | V | 707 | Unknown | Face from Web |
| GenVideo Chen et al. (2024) | V | 2.3M | 16 Open-Source Models, Pika, | General Web |
| MoonValley, Sora, MorphStudio | ||||
| AIGVDBench Ma et al. (2026a) | V | 422K | 20 Open-, 11 Closed-Source Models | General Web |
| Method | T | Approach | Year | Synthetic Training Data Type |
|---|---|---|---|---|
| C2P-CLIP Tan et al. (2025) | I | PT | 2025 | GAN + diffusion images |
| ForgeLens Chen et al. (2025) | I | PT | 2025 | GAN images |
| D3-Im Yang et al. (2025) | I | PT | 2025 | GAN + diffusion images |
| B-Free Guillaro et al. (2025) | I | TD | 2025 | Diffusion images |
| Community-Forensics Park and Owens (2025) | I | TD | 2025 | Diffusion images |
| WaveRep Corvi et al. (2025) | V | TD | 2025 | Pyramid flow videos |
| Category | Real | Grok | VEO3.1 | Sora2 | Selection |
|---|---|---|---|---|---|
| UCF-Crime | 444 | 436 | 413 | 329 | 307 |
| Abuse | 24 | 20 | 21 | 20 | 16 |
| Arrest | 20 | 20 | 18 | 10 | 9 |
| Arson | 29 | 29 | 29 | 19 | 19 |
| Assault | 39 | 37 | 36 | 25 ∗ | 21 |
| Burglary | 61 | 61 | 61 | 50 | 50 |
| Data | number | avg duration | fps | resolution | avg bitrate |
|---|---|---|---|---|---|
| Raw data | |||||
| Real data | 460 | 7.94 | 6 to 30 | 320 240 to 1920 1080 | 389.58 |
| Grok | 460 | 8.04 | 24 | 640 480 and 1280 720 | 3036.63 |
| VEO | 460 | 8.00 | 24 | 1280 720 | 2335.15 |
| Sora2 | 460 | 8.30 | 30 | 1280 720 | 4295.10 |
| Cleaned data | |||||
| Method | Grok | VEO3.1 | Sora2 | Global |
|---|---|---|---|---|
| AUC / AP / Acc | ||||
| B-Free Guillaro et al. (2025) | 0.427 / 0.473 / 0.460 | 0.359 / 0.401 / 0.359 | 0.465 / 0.467 / 0.465 | 0.417 / 0.704 / 0.450 |
| C2P-CLIP Tan et al. (2025) | 0.409 / 0.426 / 0.510 | 0.419 / 0.430 / 0.502 | 0.402 / 0.424 /0.504 | 0.410 / 0.684 / 0.505 |
| ForgeLens Chen et al. (2025) | 0.322 / 0.411 / 0.500 | 0.327 / 0.400 / 0.500 | 0.399 / 0.440 / 0.500 | 0.349 / 0.416 / 0.500 |
| D3-Im Yang et al. (2025) | 0.510 / 0.542 / 0.508 | 0.621 / 0.630 / 0.511 | 0.469 / 0.521 / 0.510 | 0.533 / 0.790 / 0.509 |
| C-Forensics Park and Owens (2025) | 0.711 / 0.710 / 0.505 | 0.693 / 0.677 / 0.502 | 0.750 / 0.735 / 0.507 | 0.718 / 0.875 / 0.504 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Base dataset | Real | Grok | VEO3.1 | Sora2 | Selection |
| Abuse | 24 | 20 | 21 | 20 | 16 |
| Arrest | 20 | 20 | 18 | 10 | 9 |
| Arson / Fire | 48 | 48 | 48 | 38 | 38 |
| Assault | 49 | 47 | 45 | 28 | 24 |
| Burglary | 61 | 61 | 61 | 50 | 50 |
| Explosion | 74 | 74 | 57 | 60 | 48 |