Organizations: Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences (CAS), Beijing, China · University of Chinese Academy of Sciences, Beijing, China · School of Advanced Interdisciplinary Sciences, University of Chinese of Sciences, Beijing, China
Interactive video generation (IVG) models have achieved remarkable progress in producing controllable visual content guided by user-defined actions, yet their security vulnerabilities remain largely unexplored. In this paper, we present the first systematic study of backdoor attacks against the interactivity of IVG models. Based on this attack surface, we propose BadAction, which leverages action-guided triggers to achieve the attack. Specifically, BadAction implants predefined motion patterns into the action sequences of backdoor samples and associates them with a static target video. Once triggered, the backdoored model generates frozen future frames that no longer respond to subsequent user actions, while preserving normal behavior on benign action sequences. In addition, we explore a stealthier attack in which multimodal triggers jointly poison action, text, and image inputs. Experiments show that BadAction achieves average attack success rates of 91.0% with action-only triggers and 80.4% with multimodal triggers. Moreover, extensive defense evaluations show that BadAction successfully bypasses existing backdoor detection methods, revealing a critical security gap in the interactive video generation pipeline. Project page: https://wsad55.github.io/badaction01/.
Figures & tables
Figure 1: Overview of the proposed BadAction attack on interactive video generation.
Figure 2: Overview of BadAction. The adversary injects a predefined action trigger, constructs poisoned action-video pairs with a static target, and fine-tunes the IVG model with a mixed objective. At inference, triggered actions generate frozen frames that ignore subsequent user actions.
Figure 3: Qualitative comparison on the action trigger. Each row shows frames sampled from one generated video: benign actions (top) keep following user actions, while triggered actions (bottom) yield identical frames after the trigger position.
Figure 4: Qualitative results of the text-image dual-modal trigger. Top : benign text and condition-frame inputs. Bottom : the trigger keyword in the text prompt and the SIG patch in the condition frame make the model generate frozen frames.
Figure 5: Qualitative results of the tri-modal trigger. Top : benign action, text, and condition-frame inputs. Bottom : the trigger motion pattern, the trigger keyword, and the SIG patch jointly activate the backdoor and freeze subsequent frames.
Modality Category
Attack Method
Trigger Modality
ASRSSIM S-T (%) ↑
ASRHuman (%) ↑
Single-Modality
BadNets ( Gu et al., 2019 )
Image
63.2
60.1
Blended ( Chen et al., 2017 )
Image
60.5
53.1
Sig ( Barni et al., 2019 )
Image
69.4
61.3
ReFool ( Liu et al., 2020 )
Image
26.4
20.7
WaNet ( Nguyen and Tran, 2021 )
Image
48.6
47.3
BadNets-T ( Gu et al., 2019 )
Text
29.7
26.9
Table 1: Attack performance of BadAction under different trigger modalities. Bold denotes the proposed BadAction method.
Modality Category
Attack Method
Trigger Modality
CLIPSIM (%) ↑
FVD ↓
Single-Modality
Benign
–
85.1
345.7
BadNets ( Gu et al., 2019 )
Image
82.5
777.8
Blended ( Chen et al., 2017 )
Image
80.3
834.2
Sig ( Barni et al., 2019 )
Image
82.4
2125.2
ReFool ( Liu et al., 2020 )
Image
77.9
1036.0
WaNet ( Nguyen and Tran, 2021 )
Image
82.2
803.8
Table 2: Generation quality of different trigger modalities. CLIPSIM measures image-text semantic alignment ( ↑ ), and FVD measures video quality ( ↓ ). Bold denotes the proposed BadAction method.
Figure 6: Ablation studies on the trigger length k , the poisoning ratio θ , the loss weighting factor λ , and the training epochs e , using both ASRSSIM S-T and ASRHuman as metrics.
Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as arrows, sketches, and emojis orchestrate complex video dynamics with unprecedented controllability. However, these seemingly innocuous static cues can be interpreted by models as executable temporal instructions, unfolding into harmful actions in the generated videos. Despite the severity of this threat, existing safety benchmarks remain predominantly focused on text-based and content-only image-based jailbreaks, leaving implicit visual prompt attacks insufficiently explored. To bridge this gap, we present VVA-Bench, the first systematic benchmark for evaluating video generation safety under categorized vision-centric prompt attacks. Extensive experiments on VVA-Bench demonstrate that state-of-the-art models are highly susceptible to such attacks, with Attack Success Rates (ASR) reaching 100.0% on Wan 2.7 and 74.8% on Veo 3.1. To mitigate these risks, we propose VPA-Guard, a retrieval-augmented and self-evolving defense framework. By leveraging few-shot reasoning to identify latent malicious intents, our method reduces the attack ASR by 44.2% and the harmfulness score by 73.4% on average, while maintaining the model's utility for legitimate user edits. Our work provides both a rigorous benchmark and an effective defense strategy to advance safe and socially responsible multimodal generation.
Yining Sun, Haoyu Kang, Jiajun Wu +7
1Tsinghua University · 3Central South University · 4South China Normal University +3
The rapid evolution of video generation has shifted the paradigm from pure text-driven to multi-conditional controllable generation, with reference images now widely adopted as conditional inputs to achieve superior spatiotemporal consistency. While these reference images serve as powerful visual anchors that significantly enhance controllability, their impact on safety remains largely unexplored. In this work, we reveal the visual anchoring effect: by enforcing consistency, the mechanism prevents the generated content from drifting away from the original harmful intent, thereby eliminating the model's natural safety escape route from harmful to benign content. Consequently, visual anchors inherently increase the safety risk---this is the price of consistency. Building on this insight, we propose Decoupling Intent via Visual Anchors (DIVA), a training-free multimodal jailbreak framework for video generation that exploits this vulnerability. DIVA decouples harmful intent into a static visual anchor image and a dynamic motion text prompt, and employs dual-criteria selection to balance attack stealthiness with semantic preservation. Extensive experiments across various leading commercial platforms and mainstream open-source video generation models demonstrate that DIVA achieves a substantially higher Attack Success Rate than existing text-only methods. To facilitate future research, we additionally contribute TI2VSafetyBench, the first safety benchmark for multi-conditional video generation.
Peng Li, Qianqian Xu, Yangbangyan Jiang +2
School of Computer Science and Technology, University of Chinese Academy of Sciences, Beijing, China · State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China · School of Artificial Intelligence and Robotics, Hunan University, Changsha, China +1
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment through voice instructions. Vidu S1 supports infinite-length real-time video generation without blurring, drift, or visual distortion. Built with TurboDiffusion and TurboServe, Vidu S1 outputs 540p real-time videos at up to 42 FPS on regular consumer GPUs. Users can upload custom images of real people, anime, and pets, and choose different voice tones for personalized experiences. Experiments show that Vidu S1 achieves the best performance across all test metrics while fully meeting real-time inference requirements. A playable online demo is available at https://vidu.com/vidu-stream.