cs.CVSep 30, 2026

BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers

Authors: Zhihang Wu, Zhongqi Wang, Jie Zhang, Fengming Gu, Shiguang Shan, Xilin Chen

Organizations: Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences (CAS), Beijing, China · University of Chinese Academy of Sciences, Beijing, China · School of Advanced Interdisciplinary Sciences, University of Chinese of Sciences, Beijing, China

Abstract

Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the interactivity of IVG models. Based on this attack surface, we propose BadAction, which leverages action-guided triggers to achieve the attack. Specifically, BadAction implants predefined motion patterns into the action sequences of backdoor samples and associates them with a static target video. Once triggered, the backdoored model generates frozen future frames that no longer respond to subsequent user actions, while preserving normal behavior on benign action sequences. In addition, we explore a stealthier attack in which multimodal triggers jointly poison action, text, and image inputs. Experiments show that BadAction achieves average attack success rates of 91.0% with action-only triggers and 80.4% with multimodal triggers. Moreover, extensive defense evaluations show that BadAction successfully bypasses existing backdoor detection methods, revealing a critical security gap in the interactive video generation pipeline. Project page: https://wsad55.github.io/badaction01/.

Figures & tables

Explore similar work

CardsList
  1. VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

    Jun 24, 2026Yining Sun, Haoyu Kang, Jiajun Wu +7Video GenerationJailbreak Attacks

  2. The Price of Consistency: Exploiting Visual Anchors for Multimodal Jailbreaking in Video Generation

    Sep 7, 2026Peng Li, Qianqian Xu, Yangbangyan Jiang +2Generative Video ModelsVideo Generation

  3. Vidu S1: A Real-Time Interactive Video Generation Model

    Jul 3, 2026Jintao Zhang, Kai Jiang, Jintao Chen +24Combustion Control