cs.AISep 27, 2026

Cross-modal Translation via Conditional Latent Denoising for Video Deepfake Detection

Authors: Xinzhe Li, Youzhi Tu, Kong Aik Lee

Organizations: Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR

Abstract

The growing threat of video deepfakes necessitates multimodal detection. Beyond serving as independent indicators of authenticity, audio and visual signals have intrinsic dependencies that also provide an essential criterion for detection. Previous methods often overlook the cross-modal correspondences, hindering information transfer between domains and leaving crucial detection cues unexplored. To address this challenge, we propose a framework called Cross-modal Translation via Conditional Latent Denoising (CTCLD) for video deepfake detection. It connects the distinct distributions of heterogeneous modalities in latent spaces, enabling smooth cross-domain information transfer to improve detection performance. We first establish a Bayesian foundation by decomposing the audio-visual joint distribution. Subsequently, CTCLD translates both modalities via bidirectional latent denoising conditioned on each other, effectively capturing subtle inconsistencies in the manipulated signals. Experimental results demonstrate that the proposed CTCLD enables comprehensive domain alignment, resulting in a robust video deepfake detection approach with competitive performance.

Figures & tables

Explore similar work

CardsList
  1. Attribution-Guided Multimodal Deepfake Detection via Cross-Modal Forensic Fingerprints

    Apr 29, 2026Wasim Ahmad, Wei Zhang, Xuerui MaoDeepfake DetectionAi-Generated Video Detection

  2. CAM-VFD: Cross-Attention Multimodal Video Forgery Detection

    May 16, 2026Hoda Osama Elkhodary, Sherin Mostafa Youssef, Marwa Elshenawy +1Deepfake DetectionAi-Generated Video Detection

  3. Ensemble Deep Learning Approaches for AI-Altered Video Detection

    Jul 8, 2026Laiba Khan, Hung-Mao Wu, Wei Lin +3Ai-Generated Video DetectionDeep Learning