Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
Organizations: University of Science and Technology of China · Independent Researcher · Zhejiang University · Hefei University of Technology
Abstract
Recent audio-video generators increasingly support joint conditioning on text, images, audio, and video. These capabilities also enable attacks that exploit cross-modal interactions or obscure harmful intent to bypass safeguards and induce harmful audio-video outputs. However, existing generation-safety benchmarks have not kept pace with these advances, providing limited coverage of multimodal input combinations and obscured attack intents. To address these gaps, we introduce Multi2AV-Safety, the first full-coverage red-team benchmark for multimodal-to-audio-video generation, comprising 11,024 attack instances across all 11 non-singleton T/I/A/V conditioning configurations, 4 attack-intent categories, and 5 harm categories. Our evaluation of recent state-of-the-art models, including four multimodal-conditioned audio-video generators and eight safety guards, reveals substantial vulnerabilities in both generation and safeguarding, with multimodal compositional risk and obscured attack-intent risk emerging as two complementary challenges. Guided by these findings, we introduce PerceptGuard, an omni-modal guard integrating compositional-risk and attack-intent supervision through structured risk perception learning. By jointly training rationale generation and safety classification, it learns shared risk representations that enable a safety head to make efficient predictions at inference without rationale decoding, while retaining the ability to generate explanations on demand. Across 34 safety benchmarks, PerceptGuard combines SOTA multimodal safety detection with highly competitive unimodal performance, strengthening input-side safeguards against multimodal attacks on omni models. In particular, it improves safeguarding against the above risks, achieving an overall recall of 86.06% on Multi2AV-Safety and outperforming GuardReasoner-Omni by 14.56%.
Figures & tables
| Input | Attack intent | Harm structure | ||||||||||
| Benchmark / Dataset | T | I | A | V | Output | Direct | Jail. | Adv. | Temp. | Full | Diluted | Composed |
| I2P ( Schramowski et al., 2023 ) | – | – | – | I | – | – | – | N/A | N/A | N/A | ||
| T2ISafety ( Li et al., 2025 ) | – | – | – | I | – | – | – | N/A | N/A | N/A | ||
| T2I-RiskyPrompt ( Zhang et al., 2026 ) | – | – | – | I | – | N/A | N/A | N/A | ||||
| JailbreakDiffBench ( Jin et al., 2025 ) | – | – | – | I/V | – | – | N/A | N/A | N/A | |||
| MPJ ( Liu et al., 2025 ) | – | – | – | I | – | – | – | N/A | N/A | N/A | ||
| Config. | Attack intent counts (#) | Generator ASR (%) | Input-side Guard Recall (%) | |||||||||
| Direct | Jail. | Adv. | Temp. | [-1pt] | [-1pt] | [-1pt] davinci | Qwen3-Omni | GR-Omni | Ours | |||
| T+I | 900 | 749 | 100 | 100 | 1,849 | 86.75 | 37.37 | 76.85 | 81.48 | 58.46 | 71.07 | 85.18 |
| T+A | 900 | 400 | 300 | 100 | 1,700 | 89.24 | 52.59 | – | 70.07 | 66.82 | 70.00 | 83.12 |
| T+V | 900 | 200 | 100 | 100 | 1,300 | 92.00 | 41.85 | 76.08 | – | 76.00 | 68.85 | 80.38 |
| I+A | 900 | 400 | 100 | 100 | 1,500 | 99.07 | 20.93 | 64.87 | 85.36 | 55.53 | 78.67 | 96.53 |
| Training objectives | Supervised outputs | 32-task F1 (%) | Omni recall (%) | ||||||
| Rationale generation | Intent recognition | Request | Response | Omni- SafetyBench | Multi2AV- Safety | ||||
| 87.44 | 74.05 | 99.35 | 71.43 | ||||||
| 88.60 | 73.39 | 99.13 | 76.24 | ||||||
| 89.01 | 72.56 | 99.37 | 82.18 | ||||||
| [0pt][0.2pt] | 88.98 | 74.83 | 99.12 | 86.06 | |||||
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Axis | Types and variations |
|---|---|
| Modality | T: existing safety datasets, adapted and refined with Claude Opus 4.6; I: Stable Diffusion 1.5 ( Rombach et al., 2022 ) / HiDream-O1-Image ( Cai et al., 2026 ) ; A: TTS by CosyVoice2 ( Du et al., 2024 ) / TTS-1 ( OpenAI, 2023 ) ; V: LTX-2 ( HaCohen et al., 2026 ) / daVinci-MagiHuman ( SII-GAIR et al., 2026 ) . |
| Conditioning | 2-modal: T+I, T+A, T+V, I+A, I+V, A+V; 3-modal: T+I+A, T+I+V, T+A+V, I+A+V; 4-modal: T+I+A+V. |
| Attack intent | Direct: explicit harmful intent; Obscured: Jailbreak, Adversarial, and Temporal intents. |
| Harm evidence | Composed Harm ( ): individually benign inputs convey harm in the complete multimodal request; Diluted Harm ( ): explicit harmful evidence is embedded in benign context; Full Harm ( ): all active inputs are harmful in isolation. |