The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at https://bruceyg.github.io/ATPO-project-page/ .
Figures & tables
Figure 1 : Real-world moderation policies require different precision–recall trade-offs: fixed-operating-point binary systems may miss harmful content and incorrectly pass it to sensitive audiences, whereas our proposed ATPO enables category-aware precision–recall control in multi-label video safety detection.
Figure 2 : Overview of Adaptive Tversky Policy Optimization (ATPO). The training pipeline consists of five stages. (1) SFT warmup: The VLM is first supervised fine-tuned using harmful video category definitions. (2) GRPO training: The model generates predictions through GRPO rollouts, which are evaluated using Adaptive Tversky Reward (ATR) and model parameters are updated. (3) Category-wise FNs and FPs from the rollouts are aggregated and tracked using EMA. (4) The FN -to- FP ratio is updated using the EMA statistics. (5) ATPO controller compares the ratio with a user-specified target ratio rc∗ and updates the Tversky coefficients αt+1,c,βt+1,c for the next training step.
#
Model
SafeWatch-Bench-Real
XD-Violence
Jaccard
Micro
Macro
Binary
Jaccard
Micro
Macro
Binary
General-purpose VLMs
1
Qwen2.5VL-7B
40.66
42.17
28.74
65.94
76.54
72.29
62.78
90.21
2
Qwen2.5VL-32B
41.66
43.60
32.11
64.63
77.69
73.28
62.71
91.85
3
Qwen3VL-8B
48.69
52.45
40.32
77.76
81.46
77.15
70.87
94.76
4
Qwen3VL-235B-A22B
55.85
62.10
56.83
80.29
82.53
78.03
72.29
95.36
Table 1: Main results on SafeWatch-Bench-Real and XD-Violence . We report multi-label metrics (Jaccard, Micro F1, Macro F1) and binary detection performance (Binary F1). Models are grouped into general-purpose VLMs, video reasoning VLMs, guardrail systems, and our ATPO-trained models with G-ATR and C-ATR. Bold and underlining indicate the best and second-best results. ∗ : further trained on the respective benchmark training sets. † : SafeWatch-8B uses full videos and clips from SafeWatch-Bench.
Training
SafeWatch-Bench-Real
SafeWatch-Bench-GenAI
Jaccard
Micro
Macro
Binary
Jaccard
Micro
Macro
Binary
Zero-shot
40.66
42.17
28.74
65.94
47.18
43.30
27.18
75.54
SFT-1ep (RL initialization)
44.33
44.73
38.47
63.53
66.34
71.40
67.90
90.20
SFT-3ep (extended SFT)
67.93
72.89
71.05
92.02
62.32
70.49
68.11
87.86
RL from the same SFT-1ep checkpoint
GRPO (EM)
63.36
67.59
63.88
92.30
62.23
68.84
64.79
89.48
Table 2: Training comparison on Qwen2.5-VL-7B. GRPO uses static rewards (Exact Match and STR), whereas ATPO-G and ATPO-C use adaptive rewards G-ATR and C-ATR, respectively. All RL methods start from the SFT-1ep checkpoint. Micro, Macro, and Binary denote F1 scores. Bold and underlining indicate the best and second-best results per column.
STR
G-ATR
α/β
Precision
Recall
P/R
ETR
r∗
Precision
Recall
P/R
ETR
1 / 5
81.24
73.21
1.11
1.58
0.2
66.33
86.55
0.77
0.31
1 / 2
79.11
69.88
1.13
1.63
0.5
76.18
80.71
0.94
0.76
1 / 1
71.93
72.62
0.99
0.97
1
75.74
78.81
0.96
0.84
2 / 1
75.03
71.55
1.05
1.20
2
75.72
69.05
1.10
1.40
5 / 1
73.48
77.50
0.95
0.80
5
78.56
67.62
1.16
1.75
Table 3: Global precision-recall control on SafeWatch-Bench-Real. Comparison between Static Tversky Reward (STR) and Global Adaptive Tversky Reward (G-ATR) under varying control ratios. We report micro Precision (P), Recall (R), their ratio (P/R), and Error Type Ratio (ETR = FN/FP).
Model
r4∗
r6∗
C4 ( Misinformation )
C6 ( Extreme )
Precision
Recall
P/R
ETR
Precision
Recall
P/R
ETR
InternVL
-
-
94.44
19.54
4.83
70.00
94.74
15.79
6.00
96.00
Qwen3-VL
-
-
95.83
22.55
4.25
79.00
96.43
20.93
4.61
102.00
C-ATR
1.0
1.0
55.28
66.67
0.83
0.62
80.31
79.07
1.02
1.08
C-ATR
5.0
0.2
70.41
67.65
1.04
1.14
31.77
94.57
0.36
0.03
Table 4: Category-level precision-recall control on SafeWatch-Bench-Real. Comparison between zero-shot large general-purpose VLMs (InternVL3.5-38B and Qwen3-VL-235B-A22B) and ATPO with C-ATR under symmetric and asymmetric target ratios for Misinformation and Extreme . We report category-wise Precision, Recall, their ratio P/R, and Error Type Ratio (ETR).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Count
Percentage (%)
SafeWatch-Bench-Real
C1 (Sexual)
188
28.92
C2 (Abuse)
94
14.46
C3 (Violence)
213
32.77
C4 (Misinformation)
102
15.69
C5 (Illegal)
114
17.54
Appendix
Table 5 : Label distribution of unsafe categories in the SafeWatch-Bench-Real test set (650 harmful videos) and XD-Violence test set (500 violent videos). Category identifiers follow the original dataset annotations. Since the task is multi-label, videos may belong to multiple categories, so counts do not sum to the total number of videos.
Figure 3 : Category co-occurrence matrices for unsafe videos in the test sets of SafeWatch-Bench-Real (left) and XD-Violence (right), showing the number of videos where each pair of categories appears together.
Figure 4 : Prompt template used during ATPO training and inference on SafeWatch-Bench. Detailed descriptions of harmful categories are omitted for brevity.
Figure 5 : Micro precision-recall curves comparing ATPO (G-ATR) with different target ratios r∗ and GRPO (STR). G-ATR with r∗=1.0 improves precision across most recall levels compared with the GRPO model, while a smaller target ratio r∗=0.2 shifts the curve toward higher recall.
Figure 6 : Evolution of Error Type Ratio (ETR=FN/FP) during ATPO training under different step sizes η , with target ratio r∗=1.0 (dashed line).
Training
SafeWatch-Bench-Real
Jaccard
Micro
Macro
Binary
Zero-shot
48.69
52.45
40.32
77.76
SFT-1ep (RL initialization)
64.70
72.70
71.37
82.35
SFT-3ep (extended SFT)
74.14
78.07
77.30
93.33
RL from the same SFT-1ep checkpoint
GRPO (EM)
59.95
62.50
56.79
84.33
Appendix
Table 6: Training comparison on Qwen3-VL-8B. GRPO uses static rewards (Exact Match and STR), whereas ATPO-G and ATPO-C use adaptive rewards G-ATR and C-ATR, respectively. All RL methods start from SFT-1ep. Micro, Macro, and Binary denote F1 scores. Bold and underlining indicate the best and second-best results per column.
Model / training
VHD11K-Video-Real
Jaccard
Micro
Macro
Binary
Zero-shot
54.20
38.49
34.74
63.60
SFT-1ep (RL initialization)
63.64
63.12
50.95
80.40
SFT-3ep (extended SFT)
57.63
59.29
49.08
76.00
RL from the same SFT-1ep checkpoint
GRPO (STR)
59.03
58.28
42.17
78.20
Appendix
Table 7: Cross-dataset evaluation of Qwen2.5-VL-7B checkpoints on VHD11K-Video-Real, without target-dataset fine-tuning. Micro, Macro, and Binary denote F1 scores. Bold and underlining indicate the best and second-best results per column.
Model / training
Jaccard
Micro
Macro
Binary
SafeWatch-8B
56.43
69.72
60.67
92.90
SFT-1ep
28.66
43.93
31.94
76.12
GRPO (STR)
63.10
73.95
52.19
99.39
ATPO-G-7B
66.11
76.97
66.37
99.70
Appendix
Table 8: Multi-label-only evaluation on SafeWatch-Bench-Real, restricted to videos with at least two ground-truth harmful labels. The SFT, GRPO, and ATPO checkpoints are based on Qwen2.5-VL-7B.
Figure 7 : Qualitative comparison between GRPO with STR and ATPO with Adaptive Tversky Rewards (G-ATR and C-ATR).
Model
r4∗
r6∗
C1
C2
P
R
P/R
ETR
P
R
P/R
ETR
InternVL
–
–
88.46
84.66
1.04
1.39
56.52
18.57
3.04
5.70
Qwen3-VL
–
–
92.31
77.01
1.20
3.58
79.25
44.68
1.77
4.73
ATPO-C
1.0
1.0
84.69
94.15
0.90
0.34
63.30
73.40
0.86
0.63
ATPO-C
5.0
0.2
87.95
77.67
1.13
2.10
62.50
69.15
0.90
0.74
Model
r4∗
r6∗
C3
C4 (targeted)
Appendix
Table 9: Complete category-level control results on SafeWatch-Bench-Real. Target ratios are modified for C4 (Misinformation) and C6 (Extreme); the remaining categories are included to examine changes beyond the targeted pair. P and R denote precision and recall.
Post-hoc threshold tuning
Training-time control
τ
SFT-1ep
GRPO (STR)
r∗
ATPO-G
Precision
Recall
Precision
Recall
Precision
Recall
0.1
75.76
36.36
72.26
73.58
0.2
66.33
86.55
0.3
77.31
33.45
72.35
73.58
0.5
76.18
80.71
0.5
78.82
32.48
72.35
73.58
1.0
75.74
78.81
0.7
79.06
30.67
72.35
73.58
2.0
75.72
69.05
Appendix
Table 10: Post-hoc threshold tuning of fixed SFT-1ep and GRPO (STR) checkpoints (left), compared with ATPO-G checkpoints trained using different target ratios r∗ (right), on SafeWatch-Bench-Real.
Initialization
XD-Violence training
XD-Violence test
Method
Examples
Jaccard
Micro
Macro
Binary
Backbone
None
0 (0%)
76.54
72.29
62.78
90.21
SafeWatch ATPO-G-7B
None
0 (0%)
80.51
76.83
64.76
94.04
SafeWatch ATPO-G-7B
ATPO
395 (10%)
84.58
81.00
68.00
95.10
Backbone
SFT + ATPO
3,950 (100%)
88.17
85.24
71.95
96.69
Appendix
Table 11: Cross-taxonomy transfer and low-resource adaptation on XD-Violence. Example counts refer only to XD-Violence training data. The final row is a separately trained full-data reference.
Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to support downstream enforcement. Traditional moderation systems often rely on fragmented black-box classifiers that are difficult to maintain and lack transparency. In this paper, we present UNIVID, a UNIfied VIsion-language model for video moDeration. Unlike standard classification models, UNIVID generates policy-aware captions that serve as an interpretable intermediate representation, enabling human-verifiable decisions and multi-task reusability. While existing open-source and commercial VLMs often suffer from safety-guardrail refusals and lack fine-grained policy alignment, we develop a specialized training data recipe that combines expert human-refined labels with synthetic data to align the model with our safety guidelines. By integrating UNIVID as the core captioner, we design a novel end-to-end video moderation system that reduces violation leakage by 42.7% and overkill rate by 37.0% relatively. Meanwhile, by replacing over 1,000 policy-specific models with a single UNIVID backbone, we recycled extensive computation resources while reducing engineering maintenance overhead. To our knowledge, this is one of the first reports of a high-efficiency captioning VLM successfully supporting industrial-scale moderation and cross-functional business.
As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful videos paired with benign queries achieve higher attack success rates than the same videos paired with explicitly harmful queries. To understand the underlying mechanism of this vulnerability, we present V-DEAL, a three-level diagnostic framework that jointly analyzes this failure across model behaviour, understanding, and internal representations. By progressively ruling out perception failure and quantifying the model's internal refusal tendency, V-DEAL provides a new diagnostic perspective for analyzing the underlying mechanism of the observed vulnerability. We tested six Video LLMs on three public benchmarks and observed that models correctly recognize harmful video content with over 81% accuracy, yet the average attack success rate still reaches 48.33% under the condition pairing harmful videos with benign queries. Hidden-state analysis further shows that visual understanding activates a weaker refusal tendency than textual understanding. Furthermore, we introduce a prompt injection intervention method that reduces attack success rates by an average of 48.24 percentage points and achieves performance comparable to prior fine-tuning-based methods, providing an effective and practical means to address such safety risks in Video LLMs.
Zhetong Zhang, Honghao Fu, Miao Xu +2
University of Queensland · University of California, Merced
Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks. However, we identify two primary limitations in existing works: 1) The multi-layered characteristics of harmful videos are overlooked. Existing benchmarks predominantly formulate evaluation as a binary classification task, failing to capture implicit or deep contextual harms. 2) Explanatory rationales are completely absent. Current frameworks measure exclusively whether a model flags a video correctly rather than explaining why, turning evaluation into a black box where models can succeed through superficial shortcuts. To address these problems, we present HarmVideoBench, a multi-layered diagnostic benchmark comprising 1,379 videos paired with 4,137 multiple-choice questions. HarmVideoBench benchmarks three hierarchical dimensions: Observable Evidence, Clip-Internal Meaning, and Beyond-Clip Reasoning, aiming to evaluate models' deep understanding beyond surface cues with carefully balanced and curated samples. We evaluate 19 leading models on HarmVideoBench to assess their multidimensional understanding of harmful videos. Moreover, we introduce BCR, a benchmark-aligned method that predicts reasoning boundaries and dynamically retrieves context only when needed. Experimental results show that BCR substantially improves the base model's performance in harmful video understanding, raising the macro average from 61.7 percent to a state-of-the-art 84.4 percent.
Jiajun Wu, Haoyu Kang, Yining Sun +13
Central South University · Tsinghua University · South China Normal University +10