Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy
Authors: Abhinav Sharma, Sai Karthik Navuluru, Wang Wei, Daksh Dangi, Xiangbo Gao, Li Li, Bo Ni, Vardhan Dongre, +16 more
Organizations: University of Massachusetts Amherst · University of Texas at Dallas · Virginia Tech · Texas A&M University · University of Southern California · Vanderbilt University · University of Illinois Urbana-Champaign · Adobe Research · Arizona State University · University of Oregon · Stanford University · Dolby Laboratories · Cisco
Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.
Figures & tables
Generation Strategy
Audio Representation
Video Representation
Alignment Enforcement
Pretraining Reuse
( Section 4.1 )
( Section 4.2 )
( Section 4.3 )
( Section 4.4 )
( Section 4.5 )
Single-Tower (§ 4.1.1 )
Dual-Tower (§ 4.1.2 )
Cascaded (§ 4.1.3 )
Unified-Token (§ 4.1.4 )
Guidance-Based (§ 4.1.5 )
Waveform (§ 4.2.4 )
Mel-Spectrogram (§ 4.2.3 )
Continuous Latent (§ 4.2.1 )
Discrete Tokens (§ 4.2.2 )
Pixel (§ 4.3.4 )
2D-VAE + Temporal (§ 4.3.2 )
3D-VAE (§ 4.3.1 )
Discrete Tokens (§ 4.3.3 )
Cross-Attention (§ 4.4.1 )
Shared Pos. Enc. (§ 4.4.2 )
Discriminator (§ 4.4.3 )
Classifier Guidance (§ 4.4.4 )
Explicit Prior (§ 4.4.5 )
From-Scratch (§ 4.5.1 )
Single Pretrained (§ 4.5.2 )
Dual Pretrained (§ 4.5.3 )
Joint Audio-Video Generation
MM-Diffusion Ruan et al. (2023)
✗
✓
✗
✗
✗
✗
✓
✗
✗
✓
✗
✗
✗
✓
✗
✗
✗
✗
✓
✗
✗
CoDi Tang et al. (2023)
✗
✓
✗
✗
✗
✗
✗
✓
✗
✗
✓
✗
✗
✓
✗
✗
✗
✗
✗
✗
✓
Seeing-and-Hearing Xing et al. (2024a)
✗
✗
✗
✗
✓
✗
✗
✓
✗
✗
✓
✗
✗
✗
✗
✗
✓
✗
✗
✗
✓
Table 1: Complementary taxonomy of joint audio-video methods along five complementary design axes: generation strategy (§ 4.1 ) captures how the two modalities are produced; audio representation (§ 4.2 ) describes the latent space in which audio is generated; video representation (§ 4.3 ) describes the latent space in which video is generated; alignment enforcement (§ 4.4 ) identifies the stage at which audio-video alignment is imposed; and pretraining reuse (§ 4.5 ) characterizes the source of model weights. A check mark (✓) indicates that a method falls under the corresponding category within each axis. For the partially closed system Wan 2.5, the axes whose design is not publicly disclosed (alignment enforcement and pretraining reuse) are left without a mark.
Symbol
Meaning
v∈RTv×H×W×3
video, Tv frames of size H×W
a∈RTa
audio signal of Ta samples
x=(v,a)
audio-visual clip
X=V×A
space of audio-visual clips
fps,sr
frame rate, audio sampling rate
τ
shared clip duration
Table 2: Notation. We summarize the symbols used throughout this work. Subscripts denote modality and primes denote edited outputs.
Task
Given
Generated
Learned Distribution
Primary Correspondence
Section
Unconditional joint
∅
v,a
pθ(v,a)
semantic + temporal
§ 6.1
Text-to-audio-video
ct
v,a
pθ(v,a∣ct)
semantic + temporal
§ 6.2
Image-to-audio-video
ci
v,a
pθ(v,a∣ci)
semantic + temporal
§ 6.3
Talking-head / speech-driven
ci,cs
v,a†
pθ(v,a∣ci,cs)
lip-sync
§ 6.4
Music-driven video
cm
v
pθ(v∣cm)
beat / rhythm
§ 7.2
Video-to-audio / Foley
v
a
pθ(a∣v)
temporal (onset)
§ 7.1
Table 3: A unified view of the tasks covered. Each task is an instance of Problems 1 through 3 : it models the joint distribution over audio-visual pairs or one of its conditionals. Given lists the observed signals, Generated the produced signals, and Primary Correspondence the form of audio-visual agreement (Def. 1 ) that dominates evaluation. The table doubles as a map from a task to the section that treats it. † For talking-head generation the model re-synthesizes the speech track: cs conditions the output and a^ is emitted synchronized to v^ , so speech appears as both condition and output.
Object with sound, character with voice, ambience addition
Adding a passing car with engine noise
Joint Removal
A+V
Object removal with sound suppression, character removal
Removing a person and their voice
Joint Replacement
A+V
Object swap, scene swap, character swap
Replacing a dog with a cat (visual + sound)
Table 4: A taxonomy of audio-visual edits: nine categories broken into 28 edit types, each with representative edits and an example use case. Each type is additionally annotated by its dominant modality coupling: A+V denotes a genuinely joint edit that requires reasoning over both modalities; V → A denotes a video-driven edit with an audio consequence (or audio derived from video); A → V denotes the reverse; V and A denote edits that are primarily single-modality, with the other modality passive or unchanged. Categories are color-coded for clarity.
Dataset
Year
Domain
Approx. Scale
Cap.
Primary Task
AudioSet Gemmeke et al. (2017)
2017
in-the-wild events
∼ 2M clips, 10s each
✗
AV pretraining, V2A
VGGSound Chen et al. (2020a)
2020
in-the-wild events
∼ 200k clips, 10s each
✗
V2A, joint generation
Kinetics Kay et al. (2017)
2017
human actions
∼ 650k clips
✗
AV pretraining
Greatest Hits Owens et al. (2016)
2016
object impacts
∼ 1k videos
✗
Foley / impacts
MUSIC Zhao et al. (2018)
2018
instrument solos/duets
714 videos
✗
music V2A, separation
URMP Li et al. (2019)
2019
classical ensembles
44 multi-track pieces
✗
music V2A, separation
Table 7: Representative datasets for joint and cross-modal audio-visual generation. We group datasets by domain and report approximate scale and whether text captions are available ( Cap. ). The rightmost column lists the task each dataset most directly supports. Scales are approximate and refer to the commonly used release.
Benchmark
Task
Property Tested
JavisBench Liu et al. (2026a)
T2AV
quality and synchronization in diverse scenes
SAVGBench Shimada et al. (2026)
joint gen.
spatial alignment between first-order-ambisonics audio and video
AVGen-Bench Zhou et al. (2026)
T2AV
aesthetics vs. semantic reliability (text, speech, physics, music)
AV-Phys Bench Cui et al. (2026)
joint gen.
physical commonsense across steady and transition scenes
Table 8: Recent benchmarks for joint audio-visual generation. Each benchmark fixes a prompt set and an evaluation protocol; the columns name the task targeted and the property tested.
Metric
Measures
Modality
Better
Per-modality quality
FID Heusel et al. (2017) , FVD Unterthiner et al. (2018)
visual fidelity (distribution distance)
V
↓
Inception Score Salimans et al. (2016)
visual quality and diversity
V
↑
FAD Kilgour et al. (2019)
audio fidelity (distribution distance)
A
↓
KL (audio classifier)
audio semantic match
A
↓
Condition alignment
Table 9: Evaluation metrics organized by what they measure. Per-modality quality metrics score one stream in isolation; condition-alignment metrics score agreement with the input c ; cross-modal alignment metrics realize the score S of Def. 1 ; and editing metrics score the two competing requirements of Problem 3 . Modality indicates the streams compared (V video, A audio, T text), and Better the preferred direction.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Tasks
Conditioning Signals
Audio Targets
Method
T2AV
V2A/Foley
A2V
Joint
Editing
Text
Video
Audio
Image/Ref.
Spatial Ctrl.
Temporal Ctrl.
SFX/Foley
Speech
Music
Ambience
Primary Domain
Visually Indicated Sounds ( Owens et al., 2016 )
✗
✓
✗
✗
✗
✗
✓
✗
✗
✗
✓
✓
✗
✗
✗
Foley / impacts
Visual to Sound ( Zhou et al., 2018 )
✗
✓
✗
✗
✗
✗
✓
✗
✗
✗
✓
✓
✗
✗
✗
in-the-wild SFX
Visually Aligned Sound ( Chen et al., 2020b )
✗
✓
✗
✗
✗
✗
✓
✗
✗
✗
✓
✓
✗
✗
✗
general V2A
Foley Music ( Gan et al., 2020 )
✗
✓
✗
✗
✗
✗
✓
✗
✗
✗
✓
✗
✗
✓
✗
video-to-music
Rhythmic Soundtracks ( Gan et al., 2021 )
✗
✓
✗
✗
✗
✗
✓
✗
✗
✗
✓
✗
✗
✓
✗
human movement
Appendix
Table 10: Unified taxonomy of video-audio generation and editing methods. tasks distinguish whether a method generates audio-video from text (T2AV), synthesizes audio for a given video (V2A/Foley), synthesizes video from audio (A2V), generates video and audio jointly (Joint), or supports editing. conditioning signals capture the user/model inputs used to steer generation. audio targets indicate what acoustic layers are explicitly modeled. A check mark indicates that the method supports the corresponding capability.
Controls
Deployment / Editing
Modeling Mechanism
Method
Text
Audio Ref.
Click/Mask
Onset/Rhythm
Long-form
Stereo/Spatial
Online
Preserve Src.
LDM
Flow/RF
DiT
AR/Masked
ControlNet
Guidance/Adapter
MLLM/CoT
Diff-Foley ( Luo et al., 2023 )
✗
✗
✗
✓
✗
✗
✗
✗
✓
✗
✗
✗
✗
✗
✗
Foley Analogies ( Du et al., 2023 )
✗
✓
✗
✓
✗
✗
✗
✗
✗
✗
✗
✗
✗
✓
✗
Seeing-and-Hearing ( Xing et al., 2024a )
✓
✓
✗
✗
✗
✗
✗
✗
✓
✗
✗
✗
✗
✓
✗
V2A-Mapper ( Wang et al., 2024a )
✗
✗
✗
✗
✗
✗
✗
✗
✗
✗
✗
✗
✗
✓
✗
Video-Foley ( Lee et al., 2025 )
✓
✓
✗
✓
✗
✗
✗
✗
✓
✗
✗
✗
✓
✗
✗
Appendix
Table 11: Fine-grained taxonomy of video-to-audio, Foley, and audio-following-video-edit methods. controls summarize how a user or upstream system steers the generated soundtrack. deployment / editing constraints distinguish long-form, stereo/spatial, online, and source-preserving settings. modeling mechanisms summarize the dominant generator or alignment mechanism.
Method
GAN/VQ
AR/Transformer
LDM
Flow/RF
Masked
AV Align.
Joint Denoise
Frozen/Guidance
MLLM/Agent
Main Mechanism
MM-Diffusion ( Ruan et al., 2023 )
✗
✗
✗
✗
✗
✗
✓
✗
✗
joint multi-modal U-Net
Diff-Foley ( Luo et al., 2023 )
✗
✗
✓
✗
✗
✓
✗
✗
✗
CAVP + latent diffusion
Seeing-and-Hearing ( Xing et al., 2024a )
✗
✗
✓
✗
✗
✗
✗
✓
✗
diffusion latent aligner
V2A-Mapper ( Wang et al., 2024a )
✗
✗
✗
✗
✗
✗
✗
✓
✗
foundation-model mapper
Video-Foley ( Lee et al., 2025 )
✗
✗
✓
✗
✗
✗
✗
✓
✗
RMS two-stage control
FoleyCrafter ( Zhang et al., 2026 )
✗
✗
✓
✗
✗
✗
✗
✓
✗
semantic adapter + temporal controller
Appendix
Table 12: Modeling-mechanism taxonomy for representative audio-video generation and editing methods. The table bridges the task-level taxonomies and method sections: it identifies whether a method primarily uses GAN/VQ, autoregression, latent diffusion, flow/rectified-flow, masked modeling, AV alignment losses, joint denoising, frozen-model guidance/adapters, or MLLM/agent-style reasoning.
Year
Method
Family
Instr.
Text
Mask/Box
Point/Traj.
Pose
Image/Style
Motion
Appearance
Inpaint
Tuning-free
Fine-tune
Attention
Latent
Cond. Branch
Canonical
2024
VIA ( Gu et al., 2024a )
Temporal adaptation
✗
✓
✓
✗
✗
✗
✓
✓
✗
✗
✓
✗
✗
✓
✗
2024
Slicedit ( Cohen et al., 2024 )
Temporal adaptation
✗
✓
✗
✗
✗
✗
✗
✓
✗
✓
✗
✗
✓
✗
✗
2024
Factorized Diffusion Distillation ( Singer et al., 2024 )
Temporal adaptation
✗
✓
✗
✗
✗
✗
✓
✓
✗
✗
✓
✗
✗
✗
✗
2024
MaskINT ( Ma et al., 2024 )
Temporal adaptation
✓
✓
✓
✗
✗
✗
✗
✓
✓
✗
✓
✗
✗
✓
✗
2023
Fairy ( Wu et al., 2024a )
Temporal adaptation
✓
✓
✗
✗
✗
✗
✗
✓
✗
✗
✓
✗
✗
✓
✗
2024
VidToMe ( Li et al., 2024b )
Temporal adaptation
✗
✓
✗
✗
✗
✗
✗
✓
✗
✓
✗
✗
✓
✗
✗
Appendix
Table 13: Video editing methods most relevant to audio-video generation pipelines: temporal-adaptation, training-modification, and conditioning-branch families. These methods edit the visual stream; audio-video systems such as CoherentAVEdit can subsequently regenerate or adapt the soundtrack to match the edited result. The columns separate the user control signal, edit target, and implementation strategy.
Year
Method
Family
Instr.
Text
Mask/Box
Point/Traj.
Pose
Image/Style
Motion
Appearance
Inpaint
Tuning-free
Fine-tune
Attention
Latent
Cond. Branch
Canonical
2024
VideoGrain ( Yang et al., 2025 )
Attention injection
✗
✓
✗
✗
✗
✗
✓
✓
✗
✓
✗
✓
✗
✗
✗
2024
AnyV2V ( Ku et al., 2024 )
Attention injection
✓
✓
✓
✗
✗
✓
✓
✓
✓
✓
✗
✓
✗
✗
✗
2024
CoCoCo ( Zi et al., 2025 )
Attention injection
✓
✓
✓
✗
✗
✗
✗
✓
✓
✓
✗
✓
✗
✗
✗
2024
Object-Centric Diffusion ( Kahatapitiya et al., 2024 )
Attention injection
✓
✓
✓
✗
✗
✗
✗
✓
✗
✓
✗
✓
✗
✗
✗
2024
UniEdit ( Bai et al., 2024 )
Attention injection
✓
✓
✗
✗
✗
✗
✓
✓
✗
✓
✗
✓
✗
✗
✗
2023
Make-A-Protagonist ( Zhao et al., 2023b )
Attention injection
✓
✓
✓
✗
✗
✓
✓
✓
✗
✗
✓
✓
✗
✗
✗
Appendix
Table 14: Video editing methods most relevant to audio-video generation pipelines: attention/latent/canonical/interactive families. This continuation covers the attention-injection, motion-feature-injection, latent-manipulation, canonical-representation, point/pose-conditioning, and human(-object)-animation families.
Year
System
T2AV
V2A
A2V
Joint
Video Edit
Speech/SFX
Music/Amb.
Notes
2024
Movie Gen ( Polyak et al., 2024 )
✓
✓
✗
✓
✓
✓
✓
research model; video, audio, personalization, editing
2024
Google V2A ( Google DeepMind, 2024 )
✗
✓
✗
✗
✓
✓
✓
research system; video pixels + optional audio prompt
Table 15: Commercial and foundation-system capabilities for video generation with audio and audio-for-video editing. These systems are included because they materially define current user-facing capabilities, even when model details or weights are not fully released.
Method
Family
Code
Repository / Weights
MM-Diffusion Ruan et al. (2023)
Joint gen.
✓
https://github.com/researchmm/MM-Diffusion
CoDi Tang et al. (2023)
Joint gen.
✓
https://github.com/microsoft/i-Code (i-Code-V3)
Seeing-and-Hearing Xing et al. (2024a)
Joint gen.
✓
https://github.com/yzxing87/Seeing-and-Hearing (V2A released; other tasks pending)
AV-DiT Wang et al. (2024b)
Joint gen.
✗
—
MM-LDM Sun et al. (2024)
Joint gen.
✗
placeholder repository only (no code or weights released)
Movie Gen Polyak et al. (2024)
Joint gen.
✗
— (benchmark data only)
Appendix
Table 16: Code and artifact availability for the methods we cover (checked July 2026). ✓: official code released, repository listed; ✗: no functional official release — repositories that exist but hold no code or weights (announcement placeholders, samples-only or dataset-only repositories) count as ✗ and are annotated. Availability changes quickly; each row was checked individually at the date above.
Visual and acoustic events in the physical world are inherently coupled, yet existing video editing methods typically adopt decoupled pipelines, lacking bidirectional modality interaction. This results in two key limitations: (i) audio-visual desynchronization and (ii) contextual conflicts between generated audio and preserved content. To address these, we propose SpongeBob, the first end-to-end audio-visual joint editing framework featuring bidirectional cross-modal interaction. For synchronization, a Sync-Aware Mechanism aligns visual edits with sound events via bidirectional attention, temporal alignment, and spatial constraints. For contextual consistency, a Context-Aware Module leverages acoustic and visual context attention to prevent semantic clashes. Additionally, we introduce Sync-Preserving Training and Guidance (SPTG) to enhance alignment without degrading quality. Due to the scarcity of paired data, we construct a scalable data pipeline and a large-scale subject-level dataset. We also propose SpongeBob-Bench for systematic evaluation. Experiments show SpongeBob significantly outperforms existing baselines, improving Sync-C by 30% and Ctx-F1 by 12.5%. Our project page is available at: https://hy-spongebob.github.io/.
Sen Liang, Cong Wang, Fengbin Guan +6
University of Science and Technology of China · 2Tencent Hunyuan
Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cross-modal synchronization from scratch. We present OmniVAE, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations. Beyond reconstruction, OmniVAE uses a segment-level audio-video contrastive objective to capture temporal-semantic correspondence and align the two latent spaces. In parallel, it distills features from pretrained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces. Extensive experiments show that both objectives consistently improve the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation. These findings underscore the importance of learning unified representations as a foundation for omnimodal modeling.1
Jun Zhan, Chen Yang, Yitian Gong +23
1Fudan University · 2Shanghai Innovation Institute · 3MOSI Intelligence +1
Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed track. This limits source-level control and differs from practical audiovisual workflows, where speech, music, sound effects, and ambient sounds are represented as separate editable tracks. We introduce Soundwich, a training-free framework that transforms a frozen joint audio-video flow-matching model into a generator of multiple synchronized, independently editable audio stems coupled to a shared video. Soundwich generates separate audio stems with explicit control over their temporal activity. To keep separately generated sounds coherent, we introduce a shared scene representation that communicates global audiovisual context across stems while preserving their source-level separation. We further route cross-modal interactions between each audio stem and its corresponding visual source, improving audiovisual consistency. The resulting stems remain synchronized with the video and can be independently retimed, muted, replaced, or remixed. Experiments and human evaluations show improved temporal control, source separation, and naturalness, while enabling flexible source-level editing within coherent audiovisual generation. Code is available at https://github.com/CodyNing/Soundwich.