cs.CLOct 1, 2026

Controllable Multi-label Video Safety Detection via Adaptive Tversky Policy Optimization

Authors: Guangyu Yang, Jingbiao Mei, Mingsheng Sun, Jinghong Chen, Yingtong Bu, Pengda Qin, Da Chen, Bill Byrne

Organizations: University of Cambridge · Xiaohongshu Inc. · AntGroup · Tencent Company, China · University of Bath

Abstract

The rapid growth of video-based social media has increased users' exposure to harmful content, creating a need for reliable automated video safety detection. Although recent Vision-Language Models (VLMs) show strong video understanding capabilities, existing harmful video detection systems face two key limitations: they typically reduce safety detection to binary classification, overlooking the inherently multi-label nature of unsafe videos, and they rely on static training objectives that do not support controllable precision-recall trade-offs, though the desired operating point may vary across moderation pipelines and unsafe categories. To address these gaps, we propose Adaptive Tversky Policy Optimization (ATPO), a reinforcement learning framework for Multi-label Video Safety Detection (Multi-VSD). ATPO introduces the Adaptive Tversky Reward (ATR), which dynamically adjusts false-positive and false-negative penalties during training to enable controllable precision-recall trade-offs. Experiments on SafeWatch-Bench and XD-Violence show that ATPO substantially improves multi-label performance, increasing the Jaccard Index from 40.66 to 75.44 on SafeWatch-Bench-Real. Moreover, ATR enables reliable steering of the precision-recall operating point, supporting deployment scenarios with heterogeneous policy requirements. Code and checkpoints are provided at https://bruceyg.github.io/ATPO-project-page/ .

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. UNIVID: Unified Vision-Language Model for Video Moderation

    Jun 4, 2026Kejuan Yang, Yizhuo Zhang, Mingyuan Du +6Video CaptioningContent Moderation

  2. V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

    Jul 23, 2026Zhetong Zhang, Honghao Fu, Miao Xu +2Attack-Success Rate

  3. HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

    Jun 25, 2026Jiajun Wu, Haoyu Kang, Yining Sun +13Long-Video BenchmarksLarge Multimodal Models