MTOR: Generalizable AI-Generated Video Detection with Multimodal Semantics and Temporal Over-Regularity
Authors: Hang Wang, Chao Shen, Lei Zhang, Zhi-Qi Cheng
Organizations: Xi’an Jiaotong University, Xi’an, China · The Hong Kong Polytechnic University, Hong Kong, China · Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA, USA
The rapid evolution of video generation has narrowed the perceptual gap between authentic and synthetic videos, making generalizable AI-generated video detection increasingly challenging. Existing detectors predominantly rely on visual representations, leaving caption-derived textual semantics underexplored. Meanwhile, temporal regularity in fine-grained visual representations has received limited attention. We find that caption-derived textual representations provide complementary discriminative cues to global visual representations. Our analysis further reveals that AI-generated videos exhibit stronger temporal persistence and lower temporal variability, a pattern we term temporal over-regularity (TOR). Based on these findings, we propose MTOR with a multimodal branch and a TOR component. The multimodal branch integrates global visual and caption-derived textual representations, while the TOR component models temporal over-regularity at three levels: coarse inter-frame continuity, fine-grained token correspondence, and frame-to-video stability. Extensive evaluations on five benchmarks covering 46 generator variants demonstrate state-of-the-art overall performance against 16 representative baselines, while robustness experiments confirm strong resilience to twelve real-world video perturbations. Code and models will be released at https://github.com/hwang-cs-ime/MTOR.
Figures & tables
Variant
Visual
Caption-derived Text
AP
AUC
ACC
Visual-only
✓
97.42
97.04
90.30
Text-only
✓
93.18
92.81
85.64
Multimodal G
✓
✓
98.81
98.66
93.80
TABLE I : Performance comparison of visual-only, text-only, and multimodal variants on GenVideo, measured by AP (%), AUC (%), and ACC (%).
Fig. 1 : Temporal over-regularity on GenVideo. (a) Three levels of temporal characterization. (b) Coarse inter-frame continuity, visualized by the temporal mean and standard deviation of adjacent-centroid similarity. (c) Fine-grained token correspondence, visualized by the temporal mean and standard deviation of adjacent-token correspondence. (d) Frame-to-video stability, visualized by the standard deviation and mean absolute temporal change of frame-to-video similarity. Each point in (b)–(d) represents one video, with real and AI-generated videos shown in blue and orange, respectively; cross markers indicate the corresponding class means. Across all three levels, AI-generated videos consistently exhibit stronger temporal persistence and reduced temporal variability than real videos.
Variant
AP
AUC
ACC
Multimodal G
97.26
97.17
89.51
TOR-only
93.22
94.59
84.19
MTOR
97.76
97.92
90.24
TABLE II : Mean performance comparison of the multimodal branch G , the TOR-only variant, and MTOR across five benchmarks, measured by AP (%), AUC (%), and ACC (%).
Fig. 2 : Overview of the proposed MTOR framework, consisting of the video-level multimodal branch, the TOR component, and source-calibrated score fusion.
Method
MS
MPS
MV
HotShot
Show-1
Gen2
Crafter
LaVie
Sora
WS
mean
FID † [ 51 ] , NeurIPS 2024
91.50
92.24
93.67
86.10
90.61
93.27
92.41
83.68
74.95
82.24
88.07
NPR † [ 52 ] , CVPR 2024
84.67
96.53
96.79
40.17
21.61
96.35
97.02
22.37
90.55
66.51
71.26
STIL † [ 18 ] , ACM MM 2021
88.21
88.68
73.07
56.91
62.07
83.96
64.87
65.54
49.86
63.60
69.68
FTCN † [ 19 ] , ICCV 2021
70.01
83.59
97.07
87.42
93.30
91.86
91.72
84.16
44.48
84.46
82.81
TALL † [ 22 ] , ICCV 2023
51.11
63.63
92.09
44.00
51.06
93.47
87.85
59.07
15.82
64.43
62.25
MINTIME † [ 23 ] , TIFS 2024
79.27
82.03
89.80
87.68
89.23
88.26
87.34
82.48
80.75
85.10
85.19
TABLE III : Comparison of AP (%) between MTOR and 16 representative baselines on GenVideo. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Method
MS
MPS
MV
HotShot
Show-1
Gen2
Crafter
LaVie
Sora
WS
mean
FID † [ 51 ] , NeurIPS 2024
90.94
91.93
93.70
85.77
91.19
93.16
92.02
82.57
73.45
81.30
87.60
NPR † [ 52 ] , CVPR 2024
93.92
99.38
99.92
26.83
18.99
98.78
99.24
42.32
97.56
76.12
75.31
STIL † [ 18 ] , ACM MM 2021
86.48
88.75
82.77
60.00
68.71
86.67
73.56
72.67
54.18
69.94
74.37
FTCN † [ 19 ] , ICCV 2021
69.76
83.25
97.18
88.69
93.45
93.33
91.62
84.70
38.30
85.74
82.60
TALL † [ 22 ] , ICCV 2023
58.46
60.54
83.24
45.21
46.24
73.30
66.37
48.40
66.36
53.84
60.20
MINTIME † [ 23 ] , TIFS 2024
81.27
85.48
90.97
90.34
90.89
88.80
89.77
83.86
79.27
85.71
86.64
TABLE IV : Comparison of AUC (%) between MTOR and 16 representative baselines on GenVideo. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Method
Floor33
Gen2
Gen2-D
HS-XL
LaVie-B
LaVie-I
Mix-SR
MS
MV
PKL
PKL-V1
Show-1
VC
ZS
mean
FID †
96.40
97.36
98.68
89.90
92.92
84.19
98.51
95.74
98.29
99.49
99.17
96.77
95.71
95.18
95.59
NPR †
99.77
99.34
99.95
47.39
76.45
72.23
99.67
98.54
99.96
99.97
99.93
69.82
99.68
98.21
90.07
STIL †
90.42
90.59
68.07
59.36
65.96
66.02
65.41
89.20
74.71
95.86
79.97
60.21
67.14
89.73
75.90
FTCN †
82.85
88.50
94.72
88.25
82.71
81.91
92.74
71.16
96.65
96.36
93.14
93.55
92.23
69.09
87.42
TALL †
63.25
70.75
77.04
46.93
52.87
52.53
78.16
62.11
83.63
65.33
70.98
48.00
60.50
51.73
63.13
MINTIME †
84.62
86.20
88.02
90.07
83.12
82.99
88.87
78.47
90.56
94.65
91.31
87.99
88.18
88.43
87.39
TABLE V : Comparison of AP (%) between MTOR and 16 representative baselines on EvalCrafter. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Method
Floor33
Gen2
Gen2-D
HS-XL
LaVie-B
LaVie-I
Mix-SR
MS
MV
PKL
PKL-V1
Show-1
VC
ZS
mean
FID †
96.26
97.46
98.58
90.05
92.32
83.42
98.23
95.27
98.28
99.48
99.24
96.79
95.26
95.29
95.42
NPR †
99.38
97.74
99.80
26.83
50.23
36.00
99.22
93.92
99.92
99.92
99.71
18.99
99.25
92.60
79.54
STIL †
90.14
91.40
79.25
61.62
73.63
71.23
74.22
87.28
84.22
97.82
88.50
66.68
74.49
90.77
80.80
FTCN †
82.78
90.76
95.11
90.29
84.51
83.00
92.31
70.99
96.65
96.73
94.01
93.62
91.27
70.82
88.06
TALL †
61.46
72.04
74.60
43.88
49.64
50.02
77.50
58.99
83.19
67.13
72.51
46.81
55.69
46.03
61.39
MINTIME †
86.88
88.01
89.41
91.71
84.07
84.23
90.16
79.60
91.68
95.08
92.33
90.08
90.16
91.29
88.91
TABLE VI : Comparison of AUC (%) between MTOR and 16 representative baselines on EvalCrafter. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Method
CVX
CVX-5B
DM
Gen-2
LaVie
OpenSora
Pika
SVD-T2I2V
VC2
ZS
mean
FID † [ 51 ] , NeurIPS 2024
93.34
91.41
97.50
98.35
96.51
87.90
99.55
95.66
96.03
90.60
94.69
NPR † [ 52 ] , CVPR 2024
81.37
81.99
99.86
99.90
63.72
88.78
99.91
99.54
60.21
78.23
85.35
STIL † [ 18 ] , ACM MM 2021
66.43
69.61
65.31
69.63
63.68
61.94
97.07
67.51
67.54
58.56
68.73
FTCN † [ 19 ] , ICCV 2021
74.24
74.83
67.98
93.99
69.16
66.95
93.94
82.81
83.39
75.64
78.29
TALL † [ 22 ] , ICCV 2023
39.59
50.72
62.36
70.78
40.40
37.30
62.69
52.62
52.66
50.66
51.98
MINTIME † [ 23 ] , TIFS 2024
85.27
85.96
77.60
84.12
79.16
90.64
92.86
79.01
83.83
86.34
84.48
TABLE VII : Comparison of AP (%) between MTOR and 16 representative baselines on VideoPhy. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Method
CVX
CVX-5B
DM
Gen-2
LaVie
OpenSora
Pika
SVD-T2I2V
VC2
ZS
mean
FID † [ 51 ] , NeurIPS 2024
93.37
91.72
97.62
98.11
96.03
86.03
99.59
95.81
95.68
89.38
94.33
NPR † [ 52 ] , CVPR 2024
72.10
73.60
99.70
99.80
42.90
83.50
99.80
99.50
47.20
52.90
77.10
STIL † [ 18 ] , ACM MM 2021
71.73
73.94
73.01
80.75
68.66
63.42
98.47
77.82
75.18
55.85
73.88
FTCN † [ 19 ] , ICCV 2021
80.91
81.01
70.20
95.21
72.88
76.37
95.90
87.07
86.53
74.88
82.10
TALL † [ 22 ] , ICCV 2023
30.05
46.36
56.54
64.80
32.94
24.82
63.65
45.40
42.48
39.44
44.65
MINTIME † [ 23 ] , TIFS 2024
85.08
87.24
82.06
86.32
76.29
89.92
93.61
80.36
85.09
89.30
85.53
TABLE VIII : Comparison of AUC (%) between MTOR and 16 representative baselines on VideoPhy. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Method
MS
OpenSora
Pika
ST2V
T2VZ
VC2
mean
FID † [ 51 ] , NeurIPS 2024
91.35
87.68
99.59
97.87
68.51
85.92
88.49
NPR † [ 52 ] , CVPR 2024
87.04
89.85
99.98
89.88
88.93
70.79
87.75
STIL † [ 18 ] , ACM MM 2021
43.68
63.15
95.28
55.46
48.68
61.49
61.29
FTCN † [ 19 ] , ICCV 2021
74.65
90.05
94.26
83.53
46.25
87.27
79.33
TALL † [ 22 ] , ICCV 2023
50.93
54.50
63.47
51.50
60.99
59.70
56.85
MINTIME † [ 23 ] , TIFS 2024
76.09
88.32
94.19
50.03
65.06
86.77
76.74
TABLE IX : Comparison of AP (%) between MTOR and 16 representative baselines on VidProM. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Method
MS
OpenSora
Pika
ST2V
T2VZ
VC2
mean
FID † [ 51 ] , NeurIPS 2024
89.86
87.36
99.62
98.10
67.01
85.65
87.93
NPR † [ 52 ] , CVPR 2024
82.61
98.56
99.84
98.92
93.32
56.70
88.33
STIL † [ 18 ] , ACM MM 2021
41.05
72.01
97.42
67.12
47.66
67.98
65.54
FTCN † [ 19 ] , ICCV 2021
70.98
90.76
95.23
87.81
50.70
88.43
80.65
TALL † [ 22 ] , ICCV 2023
45.29
56.39
66.91
50.31
57.80
53.48
55.03
MINTIME † [ 23 ] , TIFS 2024
79.40
88.99
94.81
54.21
70.92
88.96
79.55
TABLE X : Comparison of AUC (%) between MTOR and 16 representative baselines on VidProM. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Method
CogVideo
SVD
Sora
Mora
MuseV
Kling
mean
FID † [ 51 ] , NeurIPS 2024
73.10
55.95
80.06
49.13
57.29
77.29
65.47
NPR † [ 52 ] , CVPR 2024
77.88
72.85
82.56
82.13
80.27
90.58
81.05
STIL † [ 18 ] , ACM MM 2021
72.83
63.00
52.00
63.84
64.44
57.90
62.34
FTCN † [ 19 ] , ICCV 2021
85.96
87.17
44.40
74.11
97.57
74.50
77.29
TALL † [ 22 ] , ICCV 2023
55.08
59.38
51.64
55.66
58.78
57.61
56.36
MINTIME † [ 23 ] , TIFS 2024
85.70
75.72
76.62
69.53
91.05
88.57
81.20
TABLE XI : Comparison of AP (%) between MTOR and 16 representative baselines on GenVidBench. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Method
CogVideo
SVD
Sora
Mora
MuseV
Kling
mean
FID † [ 51 ] , NeurIPS 2024
76.05
53.99
77.32
40.49
54.65
76.84
63.22
NPR † [ 52 ] , CVPR 2024
74.85
65.67
77.82
75.27
73.14
88.74
75.92
STIL † [ 18 ] , ACM MM 2021
74.90
57.78
57.09
61.02
60.22
65.07
62.68
FTCN † [ 19 ] , ICCV 2021
90.15
90.87
38.91
77.56
97.98
77.41
78.81
TALL † [ 22 ] , ICCV 2023
44.12
50.88
56.63
42.29
50.09
55.51
49.92
MINTIME † [ 23 ] , TIFS 2024
85.69
78.83
76.32
74.86
91.35
89.46
82.75
TABLE XII : Comparison of AUC (%) between MTOR and 16 representative baselines on GenVidBench. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Method
GenVideo
EvalCrafter
VideoPhy
VidProM
GenVidBench
FID †
54.57
63.59
65.01
54.44
57.24
NPR †
65.41
71.36
57.00
68.04
68.39
STIL †
59.90
66.78
56.62
56.58
46.09
FTCN †
70.52
73.87
62.43
67.23
64.88
TALL †
57.47
58.42
48.78
54.02
49.23
MINTIME †
78.55
81.52
77.13
71.74
73.74
TABLE XIII : Comparison of ACC (%) between MTOR and 16 representative baselines across five benchmarks. † Results reproduced using the official code. †† Results reproduced from our implementation because no official code is available.
Variant
GenVideo
EvalCrafter
VideoPhy
VidProM
GenVidBench
Mean
Visual-only
97.42
98.06
91.59
96.58
90.24
94.78
Text-only
93.18
95.08
90.76
89.76
88.60
91.47
Multimodal G
98.81
99.11
96.61
97.11
94.66
97.26
TOR-only
94.28
96.28
95.13
89.28
91.14
93.22
MTOR
98.74
99.27
98.15
96.41
96.24
97.76
TABLE XIV : Ablation results of MTOR across five benchmarks in terms of AP (%).
Variant
GenVideo
EvalCrafter
VideoPhy
VidProM
GenVidBench
Mean
Visual-only
97.04
97.83
90.75
96.27
89.36
94.25
Text-only
92.81
95.34
92.67
88.69
89.08
91.72
Multimodal G
98.66
99.08
96.50
96.91
94.69
97.17
TOR-only
95.02
97.12
95.98
91.65
93.18
94.59
MTOR
98.64
99.30
98.33
96.61
96.71
97.92
TABLE XV : Ablation results of MTOR across five benchmarks in terms of AUC (%).
Variant
GenVideo
EvalCrafter
VideoPhy
VidProM
GenVidBench
Mean
Visual-only
90.30
91.03
83.58
89.40
82.42
87.35
Text-only
85.64
88.51
81.66
82.17
78.64
83.32
Multimodal G
93.80
95.08
87.87
88.92
81.86
89.51
TOR-only
86.14
90.03
86.07
78.59
80.14
84.19
MTOR
93.36
95.48
90.70
86.48
85.18
90.24
TABLE XVI : Ablation results of MTOR across five benchmarks in terms of ACC (%).
Caption Model
Visual & Textual Encoder
MS
MPS
MV
HotShot
Show-1
Gen2
Crafter
LaVie
Sora
WS
mean
BLIP-base
CLIP-P16
90.00
94.26
99.48
93.52
95.33
99.45
98.44
92.55
78.17
89.75
93.09
CLIP-P32
84.53
91.26
98.84
96.12
91.73
99.21
97.64
90.03
83.61
89.01
92.20
XCLIP-P16
91.12
95.32
99.40
96.96
96.29
99.32
98.58
94.54
79.92
91.49
94.29
XCLIP-P32
82.03
90.66
98.57
94.15
93.87
98.43
97.10
91.46
72.30
87.08
90.57
BLIP-large
CLIP-P16
88.18
93.16
99.31
92.11
94.15
99.28
98.01
91.52
75.76
88.50
92.00
CLIP-P32
85.63
92.35
98.80
95.14
94.52
99.22
97.65
91.21
71.29
90.56
91.64
TABLE XVII : Ablation results on GenVideo for different captioning models and visual-textual encoders in terms of AP (%).
Caption Model
Visual & Textual Encoder
MS
MPS
MV
HotShot
Show-1
Gen2
Crafter
LaVie
Sora
WS
mean
BLIP-base
CLIP-P16
90.47
93.55
99.42
93.91
95.10
99.43
98.21
91.57
78.09
88.64
92.84
CLIP-P32
86.25
90.66
98.70
96.21
92.85
99.18
97.40
89.71
84.12
88.91
92.40
XCLIP-P16
91.64
94.97
99.33
97.09
96.42
99.30
98.39
93.85
81.82
91.06
94.39
XCLIP-P32
84.26
90.45
98.39
94.88
94.16
98.29
96.76
90.89
73.82
87.19
90.91
BLIP-large
CLIP-P16
89.00
92.35
99.25
92.88
93.99
99.27
97.71
90.63
75.99
87.62
91.87
CLIP-P32
87.04
91.71
98.67
95.29
94.60
99.18
97.41
90.53
72.32
90.39
91.71
TABLE XVIII : Ablation results on GenVideo for different captioning models and visual-textual encoders in terms of AUC (%).
Caption Model
Visual & Textual Encoder
MS
MPS
MV
HotShot
Show-1
Gen2
Crafter
LaVie
Sora
WS
mean
BLIP-base
CLIP-P16
75.14
83.50
96.01
80.29
85.86
96.05
93.53
80.79
62.50
76.55
83.02
CLIP-P32
67.07
78.50
93.85
86.36
80.36
95.80
90.77
77.00
66.96
73.21
80.99
XCLIP-P16
76.00
85.29
95.85
88.00
88.21
96.16
93.28
84.71
71.43
77.81
85.67
XCLIP-P32
66.50
77.93
92.89
84.50
84.21
92.83
89.91
80.93
59.82
72.96
80.25
BLIP-large
CLIP-P16
75.21
82.86
94.57
80.29
84.86
95.14
92.49
80.68
63.39
76.10
82.56
CLIP-P32
71.71
81.36
93.29
86.43
86.57
94.64
91.17
80.50
60.71
78.75
82.51
TABLE XIX : Ablation results on GenVideo for different captioning models and visual-textual encoders in terms of ACC (%).
Fig. 3 : t-SNE visualizations of 10 generator subsets of the GenVideo test set: (a) Crafter, (b) Gen2, (c) HotShot, (d) LaVie, (e) ModelScope, (f) MoonValley, (g) MorphStudio, (h) Show-1, (i) Sora, and (j) WildScrape. Blue and orange points denote real and AI-generated videos, respectively.
Fig. 4 : Robustness comparison of MTOR and four representative baselines across 12 video perturbations using AP (%).
The proliferation of advanced AI video synthesis techniques poses an unprecedented challenge to digital video authenticity. Existing AI-generated video (AIGV) detection methods primarily focus on uni-modal or spatiotemporal artifacts, but they overlook the rich cues within the visual-textual cross-modal space, especially the temporal stability of semantic alignment. In this work, we identify a distinctive fingerprint in AIGVs, termed cross-modal temporal artifact (CMTA). Unlike real videos that exhibit natural temporal fluctuations in cross-modal alignment due to semantic variations, AIGVs display unnaturally stable semantic trajectories governed by given input prompts. To bridge this gap, we propose the CMTA framework, a cross-modal detection approach that captures these unique temporal artifacts through joint cross-modal embedding and multi-grained temporal modeling. Specifically, CMTA leverages BLIP to generate frame-level image captions and utilizes CLIP to extract corresponding visual-textual representations. A coarse-grained temporal modeling branch is then designed to characterize temporal fluctuations in cross-modal alignment with a GRU. In parallel, a fine-grained branch is constructed to capture intricate inter-frame variations from integrated visual-textual features with a Transformer encoder. Extensive experiments on 40 subsets across four large-scale datasets, including GenVideo, EvalCrafter, VideoPhy, and VidProM, validate that our approach sets a new state-of-the-art while exhibiting superior cross-generator generalization. Code and models of CMTA will be released at https://github.com/hwang-cs-ime/CMTA
Hang Wang, Chao Shen, Chenhao Lin +3
Xi’an Jiaotong University, Xi’an, China · The Hong Kong Polytechnic University, Hong Kong, China · Guangdong OPPO Mobile Communications Co., Ltd. +1
Recent advances in generative AI have democratized video creation at scale. AI-generated videos, including partially manipulated clips across visual and audio channels, pose escalating risks of semantic distortion and misuse, which motivates the need for reliable detection tools. Most existing AI-generated video detectors remain limited by single- or partial-modality of data modeling and the lack of fine-grained temporal forgery localization. To address these challenges, our primary novelty introduces a core architecture that jointly integrates an LMM semantic branch with a spatio-temporal (ST) visual branch and a multi-scale partial-spoof (PS) audio branch. This multi-modal approach enables simultaneous detection and fine-grained temporal localization of partially manipulated AI-generated video forgeries. Extensive experiments show that this approach outperforms existing state-of-the-art methods.
Dat Le, Khoa Nguyen, Xin Wang +1
Purdue University, West Lafayette · University at Albany, SUNY, New York
Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and authentic content and raising concerns about misinformation. Existing MLLM-based detectors mainly rely on supervised fine-tuning or label-level reinforcement learning, where coarse supervision limits generalization to unseen scenarios and emerging video generators. To overcome these limitations, we are the first to introduce \textbf{meta-detection} into AI-generated video detection, enabling reliable forgery detection by jointly optimizing predicted labels and supporting evidence within reinforcement learning. This paradigm requires reliable evidence signals and effective mechanisms to integrate them into label-level optimization. Textual rationales provide semantic descriptions of forgery artifacts, but their generation and verification depend on external models, making supervision vulnerable to hallucinations and semantic biases. In contrast, temporal grounding provides more objective and verifiable evidence, as manipulated intervals can be precisely controlled during forgery construction. Based on this insight, we propose an automated data construction pipeline that generates paired real-fake videos by replacing temporal segments with boundary-frame-conditioned video generation models. Furthermore, we introduce \textbf{Evidence-Guided Reward Redistribution}, which performs evidence-aware credit assignment by redistributing rewards among label-correct responses according to evidence quality. This preserves reliable label supervision while encouraging detectors to acquire fine-grained and verifiable forgery localization capabilities. Extensive experiments demonstrate that \textbf{VidForensics-M1} effectively leverages verifiable temporal evidence to achieve robust and generalizable AI-generated video detection.
Bowei Liu, Zheng Lu, Yuhan Bian +8
Tsinghua University · Peking University · Renmin University of China +1