cs.CVOct 5, 2026

BabelFake: A Multilingual Audio-Visual DeepFake Benchmark

Authors: Carlotta Segna, Joel Tschesche, Anna Rohrbach

Organizations: TU Darmstadt & Hessian.AI, Germany

Abstract

Reliable and practical audio-visual DeepFake detection requires benchmarks that reflect diverse linguistic contexts and modern data synthesis pipelines for visual as well as audio manipulations. However, existing datasets predominantly contain footage of English-speakers, often include outdated manipulation types, or overlook the audio modality. Further, many datasets feature individuals who did not consent to be used in DeepFake creation. We introduce BabelFake, a multilingual audio-visual DeepFake benchmark recorded with consenting participants. BabelFake contains 399k clips (1,323 hours) from 496 individuals spanning five languages (English, German, Italian, French, Spanish). Our modular data generation pipeline pairs 11 modern video manipulation methods with 4 voice cloning engines, distinguishing visual-only (face swapping) and joint audio-visual manipulations (lip synchronization and portrait animation). By benchmarking state-of-the-art detectors, we show that detection difficulty depends on the audio-visual generation pairing, with substantial performance degradation when authentic audio is preserved. Cross-language/demographic evaluation reveals sensitivity varying across detector architectures and training data, while human evaluation reveals that perceived realism and machine-detection difficulty do not necessarily align.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ML-ITW: A Multilingual in-the-wild Benchmark for Speech Deepfake Detection

    Mar 6, 2026Daixian Li, Jun Xue, Zhuolin Yi +4Large Audio Language ModelsMultilingual Benchmark

  2. MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

    Aug 10, 2026Yanqiu Li, Yang Xiao, Jisheng Bai +3Environmental Sound Deepfake DetectionAudio Deepfake Detection

  3. Linguistically Augmented Audio Speech Data (LinguAS)

    Jun 8, 2026Ashley R. Keaton, Zahra Khanjani, Christine Mallinson +1Audio Deepfake Detection