CCDF: A Benchmark Dataset for Deepfake Detection in Real-World Surveillance Footage
Authors: Baptiste Chopin, Thomas Swearingen, Arun Ross, Antitza Dantcheva, Christian Rathgeb
Organizations: da/sec – Biometrics and Security Research Group, Hochschule Darmstadt · Computer Science and Engineering Department, Michigan State University · STARS team, Inria Center at Université Côte d’Azur
Due to rapid advances in Generative AI, commercial video generation tools can be used to produce fabricated surveillance footage that can fool both human viewers and automated synthetic video detectors. Since these tools are so widely accessible, a malicious user can create a harmful video clip at minimal cost. The production and dissemination of such videos in high-stakes settings, such as crime reporting and elections, can misdirect emergency response efforts or distort political discourse. Existing deepfake video datasets, used by the research community to develop deepfake detection algorithms, exhibit two limitations: (1) they emphasize benign web content rather than footage of possibly malicious activity, and (2) they rely on older or open-source generators that do not represent recent advances in generative systems. We assemble CCtv DeepFakes (CCDF), a video deepfake dataset, to address both gaps. CCDF contains 1840 videos (460 real and 1380 generated) spanning 16 crime and accident categories, with generated content produced using three leading commercial systems: Grok Imagine, Google VEO 3.1, and OpenAI Sora 2. CCDF is a highly realistic, small-scale, manually annotated dataset targeting evaluation of detection models. We release three versions of the dataset: the raw generated data, a cleaned version in which video metadata are standardized between real and synthetic samples to prevent detectors from exploiting trivial cues, and an altered version simulating low-effort post-processing attacks. We evaluate CCDF with ten recent state-of-the-art detectors covering different detection approaches. Our results suggest that these approaches do not reliably distinguish CCDF's generated videos from real ones, despite their strong reported performance on existing datasets. These results further confirm that existing datasets are not well-suited to evaluating certain threats.
Figures & tables
Figure 1
Figure 1 : Overview of dataset creation: Based on a real video a prompt describing the content of the first frame as well as the entire video is extracted. Then an image model generates the first frame while a video model later generates the entire video based on the text prompt and the initial frame image. Note: the prompt shown in this figure is only an excerpt from the real prompt.
Dataset
T
# Samples
Generators
Content Style
DFDC Dolhansky et al. (2020)
V
128K
8 Open-Source Models
Face
DF40 Yan et al. (2024)
I/V
> 1M
40 Open-Source Models
Face
WildDeepfake Zi et al. (2020)
V
707
Unknown
Face from Web
GenVideo Chen et al. (2024)
V
2.3M
16 Open-Source Models, Pika,
General Web
MoonValley, Sora, MorphStudio
AIGVDBench Ma et al. (2026a)
V
422K
20 Open-, 11 Closed-Source Models
General Web
Table 1 : Comparison of select deepfake datasets. The “T” column describes the multimedia type, video (V) or image (I).
Method
T
Approach
Year
Synthetic Training Data Type
C2P-CLIP Tan et al. (2025)
I
PT
2025
GAN + diffusion images
ForgeLens Chen et al. (2025)
I
PT
2025
GAN images
D3-Im Yang et al. (2025)
I
PT
2025
GAN + diffusion images
B-Free Guillaro et al. (2025)
I
TD
2025
Diffusion images
Community-Forensics Park and Owens (2025)
I
TD
2025
Diffusion images
WaveRep Corvi et al. (2025)
V
TD
2025
Pyramid flow videos
Table 2 : Overview of the nine detectors used to benchmark CCDF. The “T” column describes the input type: image (I) or video (V), while the “Approach” column describes the broad categories: pre-trained (PT), training data (TD), or dynamic video features (DV).
Department of Computer Science, University of Bucharest, Romania · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), Abu Dhabi, UAE · Department of Computer Science, University of Central Florida, Orlando, US
School of Cyber Science and Technology, Shandong University, Qingdao, Shandong, China · Institute for Network Sciences and Cyberspace, BNRist, Tsinghua University, Beijing, China · School of Cryptologic Science and Engineering, Shandong University, Jinan, Shandong, China