cs.CVOct 5, 2026

MTOR: Generalizable AI-Generated Video Detection with Multimodal Semantics and Temporal Over-Regularity

Authors: Hang Wang, Chao Shen, Lei Zhang, Zhi-Qi Cheng

Organizations: Xi’an Jiaotong University, Xi’an, China · The Hong Kong Polytechnic University, Hong Kong, China · Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA, USA

Abstract

The rapid evolution of video generation has narrowed the perceptual gap between authentic and synthetic videos, making generalizable AI-generated video detection increasingly challenging. Existing detectors predominantly rely on visual representations, leaving caption-derived textual semantics underexplored. Meanwhile, temporal regularity in fine-grained visual representations has received limited attention. We find that caption-derived textual representations provide complementary discriminative cues to global visual representations. Our analysis further reveals that AI-generated videos exhibit stronger temporal persistence and lower temporal variability, a pattern we term temporal over-regularity (TOR). Based on these findings, we propose MTOR with a multimodal branch and a TOR component. The multimodal branch integrates global visual and caption-derived textual representations, while the TOR component models temporal over-regularity at three levels: coarse inter-frame continuity, fine-grained token correspondence, and frame-to-video stability. Extensive evaluations on five benchmarks covering 46 generator variants demonstrate state-of-the-art overall performance against 16 representative baselines, while robustness experiments confirm strong resilience to twelve real-world video perturbations. Code and models will be released at https://github.com/hwang-cs-ime/MTOR.

Figures & tables

Explore similar work

CardsList
  1. CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection

    May 1, 2026Hang Wang, Chao Shen, Chenhao Lin +3Ai-Generated Video DetectionAuthenticity

  2. Towards multi-modal forgery representation learning for AI-generated video detection and localization

    May 8, 2026Dat Le, Khoa Nguyen, Xin Wang +1Ai-Generated Video DetectionFine-Grained Video Understanding

  3. VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

    Aug 11, 2026Bowei Liu, Zheng Lu, Yuhan Bian +8Ai-Generated Video DetectionForgeries