Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy
Organizations: University of Massachusetts Amherst · University of Texas at Dallas · Virginia Tech · Texas A&M University · University of Southern California · Vanderbilt University · University of Illinois Urbana-Champaign · Adobe Research · Arizona State University · University of Oregon · Stanford University · Dolby Laboratories · Cisco
Abstract
Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.
Figures & tables
| Generation Strategy | Audio Representation | Video Representation | Alignment Enforcement | Pretraining Reuse | |||||||||||||||||||
| ( Section 4.1 ) | ( Section 4.2 ) | ( Section 4.3 ) | ( Section 4.4 ) | ( Section 4.5 ) | |||||||||||||||||||
| Single-Tower (§ 4.1.1 ) | Dual-Tower (§ 4.1.2 ) | Cascaded (§ 4.1.3 ) | Unified-Token (§ 4.1.4 ) | Guidance-Based (§ 4.1.5 ) | Waveform (§ 4.2.4 ) | Mel-Spectrogram (§ 4.2.3 ) | Continuous Latent (§ 4.2.1 ) | Discrete Tokens (§ 4.2.2 ) | Pixel (§ 4.3.4 ) | 2D-VAE + Temporal (§ 4.3.2 ) | 3D-VAE (§ 4.3.1 ) | Discrete Tokens (§ 4.3.3 ) | Cross-Attention (§ 4.4.1 ) | Shared Pos. Enc. (§ 4.4.2 ) | Discriminator (§ 4.4.3 ) | Classifier Guidance (§ 4.4.4 ) | Explicit Prior (§ 4.4.5 ) | From-Scratch (§ 4.5.1 ) | Single Pretrained (§ 4.5.2 ) | Dual Pretrained (§ 4.5.3 ) | |||
| Joint Audio-Video Generation | |||||||||||||||||||||||
| MM-Diffusion Ruan et al. (2023) | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ||
| CoDi Tang et al. (2023) | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ||
| Seeing-and-Hearing Xing et al. (2024a) | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ||
| Symbol | Meaning |
| video, frames of size | |
| audio signal of samples | |
| audio-visual clip | |
| space of audio-visual clips | |
| frame rate, audio sampling rate | |
| shared clip duration |
| Task | Given | Generated | Learned Distribution | Primary Correspondence | Section |
| Unconditional joint | semantic temporal | § 6.1 | |||
| Text-to-audio-video | semantic temporal | § 6.2 | |||
| Image-to-audio-video | semantic temporal | § 6.3 | |||
| Talking-head / speech-driven | lip-sync | § 6.4 | |||
| Music-driven video | beat / rhythm | § 7.2 | |||
| Video-to-audio / Foley | temporal (onset) | § 7.1 |
| Category | Edit/Gen. Type | Modality | Representative Edits | Example Use Case |
| Synchronization | Lip Sync | A+V | Lip sync correction, dubbing alignment, viseme generation | Aligning dubbed audio to mouth motion |
| AV Alignment | A+V | Foley alignment, beat alignment, event sync | Matching footsteps to visual steps | |
| Re-timing | A+V | Joint time stretch, slow motion, speed ramping | Slowing a scene with pitch-preserved audio | |
| Joint Content | Joint Insertion | A+V | Object with sound, character with voice, ambience addition | Adding a passing car with engine noise |
| Joint Removal | A+V | Object removal with sound suppression, character removal | Removing a person and their voice | |
| Joint Replacement | A+V | Object swap, scene swap, character swap | Replacing a dog with a cat (visual + sound) |
| Dataset | Year | Domain | Approx. Scale | Cap. | Primary Task |
| AudioSet Gemmeke et al. (2017) | 2017 | in-the-wild events | 2M clips, 10s each | ✗ | AV pretraining, V2A |
| VGGSound Chen et al. (2020a) | 2020 | in-the-wild events | 200k clips, 10s each | ✗ | V2A, joint generation |
| Kinetics Kay et al. (2017) | 2017 | human actions | 650k clips | ✗ | AV pretraining |
| Greatest Hits Owens et al. (2016) | 2016 | object impacts | 1k videos | ✗ | Foley / impacts |
| MUSIC Zhao et al. (2018) | 2018 | instrument solos/duets | 714 videos | ✗ | music V2A, separation |
| URMP Li et al. (2019) | 2019 | classical ensembles | 44 multi-track pieces | ✗ | music V2A, separation |
| Benchmark | Task | Property Tested |
| JavisBench Liu et al. (2026a) | T2AV | quality and synchronization in diverse scenes |
| SAVGBench Shimada et al. (2026) | joint gen. | spatial alignment between first-order-ambisonics audio and video |
| AVGen-Bench Zhou et al. (2026) | T2AV | aesthetics vs. semantic reliability (text, speech, physics, music) |
| AV-Phys Bench Cui et al. (2026) | joint gen. | physical commonsense across steady and transition scenes |
| Metric | Measures | Modality | Better |
| Per-modality quality | |||
| FID Heusel et al. (2017) , FVD Unterthiner et al. (2018) | visual fidelity (distribution distance) | V | |
| Inception Score Salimans et al. (2016) | visual quality and diversity | V | |
| FAD Kilgour et al. (2019) | audio fidelity (distribution distance) | A | |
| KL (audio classifier) | audio semantic match | A | |
| Condition alignment | |||
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Tasks | Conditioning Signals | Audio Targets | ||||||||||||||
| Method | T2AV | V2A/Foley | A2V | Joint | Editing | Text | Video | Audio | Image/Ref. | Spatial Ctrl. | Temporal Ctrl. | SFX/Foley | Speech | Music | Ambience | Primary Domain |
| Visually Indicated Sounds ( Owens et al., 2016 ) | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | Foley / impacts |
| Visual to Sound ( Zhou et al., 2018 ) | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | in-the-wild SFX |
| Visually Aligned Sound ( Chen et al., 2020b ) | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | general V2A |
| Foley Music ( Gan et al., 2020 ) | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | video-to-music |
| Rhythmic Soundtracks ( Gan et al., 2021 ) | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | human movement |
| Controls | Deployment / Editing | Modeling Mechanism | |||||||||||||
| Method | Text | Audio Ref. | Click/Mask | Onset/Rhythm | Long-form | Stereo/Spatial | Online | Preserve Src. | LDM | Flow/RF | DiT | AR/Masked | ControlNet | Guidance/Adapter | MLLM/CoT |
| Diff-Foley ( Luo et al., 2023 ) | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| Foley Analogies ( Du et al., 2023 ) | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ |
| Seeing-and-Hearing ( Xing et al., 2024a ) | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ |
| V2A-Mapper ( Wang et al., 2024a ) | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ |
| Video-Foley ( Lee et al., 2025 ) | ✓ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ |
| Method | GAN/VQ | AR/Transformer | LDM | Flow/RF | Masked | AV Align. | Joint Denoise | Frozen/Guidance | MLLM/Agent | Main Mechanism |
| MM-Diffusion ( Ruan et al., 2023 ) | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | joint multi-modal U-Net |
| Diff-Foley ( Luo et al., 2023 ) | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | CAVP + latent diffusion |
| Seeing-and-Hearing ( Xing et al., 2024a ) | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | diffusion latent aligner |
| V2A-Mapper ( Wang et al., 2024a ) | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | foundation-model mapper |
| Video-Foley ( Lee et al., 2025 ) | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | RMS two-stage control |
| FoleyCrafter ( Zhang et al., 2026 ) | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | semantic adapter + temporal controller |
| Year | Method | Family | Instr. | Text | Mask/Box | Point/Traj. | Pose | Image/Style | Motion | Appearance | Inpaint | Tuning-free | Fine-tune | Attention | Latent | Cond. Branch | Canonical |
| 2024 | VIA ( Gu et al., 2024a ) | Temporal adaptation | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ |
| 2024 | Slicedit ( Cohen et al., 2024 ) | Temporal adaptation | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ |
| 2024 | Factorized Diffusion Distillation ( Singer et al., 2024 ) | Temporal adaptation | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ |
| 2024 | MaskINT ( Ma et al., 2024 ) | Temporal adaptation | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ |
| 2023 | Fairy ( Wu et al., 2024a ) | Temporal adaptation | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ |
| 2024 | VidToMe ( Li et al., 2024b ) | Temporal adaptation | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ |
| Year | Method | Family | Instr. | Text | Mask/Box | Point/Traj. | Pose | Image/Style | Motion | Appearance | Inpaint | Tuning-free | Fine-tune | Attention | Latent | Cond. Branch | Canonical |
| 2024 | VideoGrain ( Yang et al., 2025 ) | Attention injection | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ |
| 2024 | AnyV2V ( Ku et al., 2024 ) | Attention injection | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ |
| 2024 | CoCoCo ( Zi et al., 2025 ) | Attention injection | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ |
| 2024 | Object-Centric Diffusion ( Kahatapitiya et al., 2024 ) | Attention injection | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ |
| 2024 | UniEdit ( Bai et al., 2024 ) | Attention injection | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✓ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ |
| 2023 | Make-A-Protagonist ( Zhao et al., 2023b ) | Attention injection | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ |
| Year | System | T2AV | V2A | A2V | Joint | Video Edit | Speech/SFX | Music/Amb. | Notes |
| 2024 | Movie Gen ( Polyak et al., 2024 ) | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | research model; video, audio, personalization, editing |
| 2024 | Google V2A ( Google DeepMind, 2024 ) | ✗ | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ | research system; video pixels + optional audio prompt |
| 2025 | Veo 3/3.1 ( Google DeepMind, 2025 ) | ✓ | ✓ | ✗ | ✓ | ✓ | ✓ | ✓ | commercial/product; native audio |
| 2025 | Sora 2 ( OpenAI, 2025 ) | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | commercial/product; synchronized dialogue/SFX |
| 2025 | Adobe Firefly Audio ( Adobe, 2025 ) | ✗ | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ | commercial creative tools; sound effects, soundtrack, speech |
| 2026 | Seedance 2.0 ( ByteDance Seed, 2026 ) | ✓ | ✗ | ✗ | ✓ | ✓ | ✓ | ✓ | commercial/product; text/image/video/audio prompts |
| Method | Family | Code | Repository / Weights |
| MM-Diffusion Ruan et al. (2023) | Joint gen. | ✓ | https://github.com/researchmm/MM-Diffusion |
| CoDi Tang et al. (2023) | Joint gen. | ✓ | https://github.com/microsoft/i-Code (i-Code-V3) |
| Seeing-and-Hearing Xing et al. (2024a) | Joint gen. | ✓ | https://github.com/yzxing87/Seeing-and-Hearing (V2A released; other tasks pending) |
| AV-DiT Wang et al. (2024b) | Joint gen. | ✗ | — |
| MM-LDM Sun et al. (2024) | Joint gen. | ✗ | placeholder repository only (no code or weights released) |
| Movie Gen Polyak et al. (2024) | Joint gen. | ✗ | — (benchmark data only) |