mAVE: A Watermark for Joint Audio-Visual Generation Models
Organizations: School of Software, Tsinghua University
Abstract
Watermarking joint audio-visual generation supports vendor copyright protection and content provenance. However, independently valid audio and video watermarks do not establish a shared generation session. An adversary can splice watermarked modalities from different sessions, causing the pair to be mistaken for the vendor's original joint output. We introduce mAVE (Manifold Audio-Visual Entanglement), a training-free watermarking framework that strengthens vendor attribution through session binding in native joint audio-visual diffusion transformers. mAVE separates public record retrieval from secret session authentication: a fixed public index locates the server record, while a randomized payload binds audio bits to a session-keyed video grid through a cryptographic digest. One prompt-conditioned joint inversion supports provider-assisted verification of both modalities against a session record, without modifying generator weights or training auxiliary watermark networks. Our analysis establishes implementation-matched distribution preservation and a full-initialization routing/clipping budget, alongside adaptive session-pool security and stable local-perturbation bounds. Experiments on LTX-2 and MOVA show comparable generation quality. mAVE achieves 99.8% true-positive rate and 0% observed false-positive rate in the evaluated swap test, and retains 99.2% true-positive rate under FrameAvg temporal averaging. Same-prompt and similarity-selected swaps further test session authentication beyond perceptual compatibility.
Figures & tables
| mAVE | Video BA | Audio BA / score | ||||||
|---|---|---|---|---|---|---|---|---|
| Model / task | Det. V/A | VS | VSeal | mAVE | AS | WavMark | Timbre | mAVE |
| LTX-2 / T2AV | 100/100 | .953 | 1.000 | .936 | 1.000 | .994 | .993 | .915 |
| LTX-2 / I2AV | 100/99.8 | .946 | .995 | .934 | 1.000 | .995 | 1.000 | .917 |
| MOVA-720p / TI2AV | 100/99.9 | .952 | .998 | .949 | 1.000 | .993 | .999 | .928 |
| Method | Sub | Mot | Dyn | Back | Img | Avg | CLAP | ImBind | Sync |
|---|---|---|---|---|---|---|---|---|---|
| Clean (No WM) | 0.983 | 0.988 | 0.715 | 0.966 | 0.524 | 0.835 | 0.442 | 0.117 | 0.965 |
| Uncoupled (Base) | 0.994 | 0.988 | 0.684 | 0.982 | 0.454 | 0.820 | 0.425 | 0.098 | 0.965 |
| mAVE (Ours) | 0.998 | 0.991 | 0.669 | 0.985 | 0.529 | 0.834 | 0.426 | 0.132 | 0.966 |
| Method | WM det. | TPR (Auth) | FNR (Auth) | TNR (Swap) | FPR (Swap) | Acc |
|---|---|---|---|---|---|---|
| Weak Base (Uncoupled) | 100% | 100% | 0% | 0% | 100% | 50.0% |
| Uncoupled + SyncNet | 100% | 96.2% | 3.8% | 76.2% | 23.8% | 86.2% |
| Uncoupled + ImageBind/CLAP | 100% | 97.8% | 2.2% | 65.4% | 34.6% | 81.6% |
| Shared Session-ID | 99.8% | 98.9% | 1.1% | 88.7% | 11.3% | 93.8% |
| SR-Rand | 99.7% | 99.7% | 0.3% | 93.2% | 6.8% | 96.5% |
| SR-ID | 99.7% | 99.7% | 0.3% | 95.7% | 4.3% | 97.7% |
| Authentic | Swap FPR (%) | |||
| Method / control | TPR (%) | Cross-prompt | Same-prompt | Selected |
| Uncoupled + SyncNet | 96.2 | 23.8 | 29.5 | 33.1 |
| Uncoupled + ImageBind/CLAP | 97.8 | 34.6 | 33.7 | 41.8 |
| Shared Session-ID | 98.9 | 11.3 | 11.2 | 10.8 |
| mAVE: no binding gate | 99.8 | 0.1 | 2.4 | 3.8 |
| mAVE: no hash or binding gate | 99.8 | 0 | 6.7 | 5.9 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| 16 | 32 | 64 | 128 | 256 | |
|---|---|---|---|---|---|
| Hoeffding bound |
| Bit-change budget | Radius | |
|---|---|---|
| 0 | 25 | |
| 4 | 29 | |
| 8 | 33 | |
| 16 | 41 |
| Framework / Model | Task | Inversion Steps ( ) | ||||
|---|---|---|---|---|---|---|
| 50 | 25 | 10 | 5 | 1 | ||
| Baseline Paradigms (Full Integration to ) | ||||||
| DDIM (ModelScope) | T2V | 1.000 | 1.000 | 1.000 | - | - |
| DDIM (SVD) | I2V | 0.999 | 0.999 | 0.999 | - | - |
| Rectified Flow (CIFAR-10) | Uncond | 1.000 | 1.000 | 1.000 | - | - |
| mAVE with Truncation (MOVA-720p) | ||||||
| Decoding Method | Video BA | Audio BA |
|---|---|---|
| Zero-Thresholding | 0.936 | 0.915 |
| Median Thresholding | 0.940 | 0.913 |
| Extraction Protocol | Video BA | Audio BA | Time (s/item) |
|---|---|---|---|
| Joint Inversion (ours) | 0.936 | 0.915 | 2.07 |
| Separate Inversion | 0.936 | 0.915 | 4.16 |
| Method | Video Extraction | Audio Extraction | Binding Check | Total Passes |
|---|---|---|---|---|
| VideoShield + AudioSeal | ODE inversion ( steps) | AudioSeal encoder (forward) | — | 2 |
| VideoShield + WavMark | ODE inversion ( steps) | WavMark encoder (forward) | — | 2 |
| mAVE (ours) | Joint ODE inversion ( steps) | 1 | ||
| Model / task | Exact index recovery (%) |
|---|---|
| LTX-2 / T2AV | 100 |
| LTX-2 / I2AV | 100 |
| MOVA-720p / TI2AV | 100 |
| Configuration | FrameAvg | FrameSwap |
|---|---|---|
| VideoShield (ModelScope) | 0.995 | 0.732 |
| VideoShield (LTX-2 T2AV) | 0.948 | 0.699 |
| mAVE (LTX-2 T2AV) | 0.927 | 0.710 |
| Category | Example Prompts |
|---|---|
| Animal | “A panda standing on a surfboard in the ocean in sunset”; “A jellyfish floating through the ocean, with bioluminescent tentacles” |
| Architecture | ‘A 3D model of a 1800s victorian house”; “Asian garden and medieval castle” |
| Food | “Slow motion cropped closeup of roasted coffee beans falling into an empty bowl”; “A serving of pumpkin dish in a plate” |
| Human | “Vincent van Gogh is painting in the room”; “An astronaut feeding ducks on a sunny afternoon, reflection from the water” |
| Lifestyle | “This is how I do makeup in the morning”; “Home deco with lighted” |
| Plant | “Yellow flowers swing in the wind”; “An elephant spraying itself with water using its trunk to cool down” |
| Method | Embedding and training | Verification objective |
|---|---|---|
| LVMark ( Jang et al., 2024 ) | Learned latent-decoder modulation and watermark decoder | Video watermark-message recovery for attribution |
| VIDSIG ( Huang et al., 2025 ) | Partial latent-decoder fine-tuning with temporal alignment | Video signature detection and recovery |
| FlowMark ( Asnani et al., 2026 ) | Trainable mask-guided media embedding | Video message recovery and integrity protection |
| V 2 A-Mark ( Zhang et al., 2024 ) | Learned video/audio watermarking modules | Copyright recovery and manipulation localization |
| Kim et al. (2025) | Trainable invertible cross-modal hiding | Audio recovery and audiovisual tamper localization |
| mAVE | Joint latent initialization; fixed generator weights; no watermark-specific training | Same-session AV pair authentication via joint inversion and a server record |
| Symbol | Meaning |
|---|---|
| Sessions, keys, and operators | |
| Distinct generation sessions, each corresponding to one inference request. | |
| Public session-record index and its decoded estimate; repeated at fixed, unencrypted video positions. | |
| Server-held session secret, with ; never embedded. | |
| Bit lengths of the public index and server secret. | |
| Fixed public index coordinates, copies of index bit , and copies per index bit. | |