BadAction: Backdoor Attacks on Interactive Video Generation via Action-Guided Triggers
Organizations: Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences (CAS), Beijing, China · University of Chinese Academy of Sciences, Beijing, China · School of Advanced Interdisciplinary Sciences, University of Chinese of Sciences, Beijing, China
Abstract
Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the interactivity of IVG models. Based on this attack surface, we propose BadAction, which leverages action-guided triggers to achieve the attack. Specifically, BadAction implants predefined motion patterns into the action sequences of backdoor samples and associates them with a static target video. Once triggered, the backdoored model generates frozen future frames that no longer respond to subsequent user actions, while preserving normal behavior on benign action sequences. In addition, we explore a stealthier attack in which multimodal triggers jointly poison action, text, and image inputs. Experiments show that BadAction achieves average attack success rates of 91.0% with action-only triggers and 80.4% with multimodal triggers. Moreover, extensive defense evaluations show that BadAction successfully bypasses existing backdoor detection methods, revealing a critical security gap in the interactive video generation pipeline. Project page: https://wsad55.github.io/badaction01/.
Figures & tables
| Modality Category | Attack Method | Trigger Modality | (%) | (%) |
|---|---|---|---|---|
| Single-Modality | BadNets ( Gu et al., 2019 ) | Image | 63.2 | 60.1 |
| Blended ( Chen et al., 2017 ) | Image | 60.5 | 53.1 | |
| Sig ( Barni et al., 2019 ) | Image | 69.4 | 61.3 | |
| ReFool ( Liu et al., 2020 ) | Image | 26.4 | 20.7 | |
| WaNet ( Nguyen and Tran, 2021 ) | Image | 48.6 | 47.3 | |
| BadNets-T ( Gu et al., 2019 ) | Text | 29.7 | 26.9 |
| Modality Category | Attack Method | Trigger Modality | CLIPSIM (%) | FVD |
|---|---|---|---|---|
| Single-Modality | Benign | – | 85.1 | 345.7 |
| BadNets ( Gu et al., 2019 ) | Image | 82.5 | 777.8 | |
| Blended ( Chen et al., 2017 ) | Image | 80.3 | 834.2 | |
| Sig ( Barni et al., 2019 ) | Image | 82.4 | 2125.2 | |
| ReFool ( Liu et al., 2020 ) | Image | 77.9 | 1036.0 | |
| WaNet ( Nguyen and Tran, 2021 ) | Image | 82.2 | 803.8 |