Multimedia editing requires generative models to interpret complex, compositional instructions and modify only the desired elements across modalities, a capability largely beyond existing systems. Current approaches are typically limited to simple, single-attribute edits within one modality, since diverse, high-quality editing pairs are hard to obtain, especially for cross-modal tasks where aligned audiovisual (AV) data is scarce. We present CrossEdit, a unified omni-modal editing model for images, audio, and video that performs AV movie scene edits zero-shot through cross-modal transfer. Our key observation is that complex instruction-following, once learned in any modality, generalizes to others. We exploit the additive nature of acoustic signals to procedurally generate large-scale audio editing pairs with compositional instructions, and show that fine-tuning on this synthetic data and self-supervised AV masked reconstruction, alongside a targeted set of curated cross-modal tasks, induces instruction-following that transfers zero-shot to unseen modality and instruction combinations. To evaluate this capability, we release CrossEditBench, a human-annotated benchmark of AV edits on movie scenes, and propose AV-FES, a metric that jointly scores instruction following and consistency. We show that our proposed techniques improve performance on zero-shot AV movie scene editing, while maintaining or improving performance on lip-synced speech editing and standard image, video, and audio editing benchmarks. Demos are available at https://wanchichen.github.io/crossedit/.
Figures & tables
Figure 1: Instruction types in CrossEditBench that CrossEdit must handle zero-shot. Models must implicitly determine which modalities to edit and what to synchronize across audio and video.
Task
Example Instruction
Pre-train
SFT
Eval.
Editing
General image editing (I2I)
“Make her hair blue”
✓
✓
§ 5.3
General video editing (V2V)
“Turn the car red”
✓
✓
§ 5.2
Speech editing † (A2A)
“Change the speech to ‘See you soon”’
✓
✓
–
Complex audio editing (A2A)
“Remove the music; make the rain louder”
✗
✓
§ D
Lip-synced speech editing (AV2AV)
“Change the speech to ‘We made it out”’
✓
✓
§ 5.1
Table 1: Training and Evaluation tasks for CrossEdit. Unlike trained lip-synced speech editing, zero-shot speech + speaker edits also change who is speaking. † Voice cloning only.
Figure 2: Architecture of CrossEdit when encoding AV data. For simplicity, we omit images, which are treated as a special case of single-frame video.
Figure 3: Editing data sources. (a) Our pipeline for mining lip-synced speech editing pairs. (b) Random audio mixtures and (c) Synthetic audio scenes provide audio editing pairs.
Figure 4: AV masked reconstruction schemes. All use the instruction “Inpaint this video”.
Overall
Visual
Audio
Alignment
subscores
subscores
Type
AV-FES ↑
V-FES ↑
Gain
Ins.
Cons.
A-FES ↑
Gain
Ins.
Cons.
ImageBind ↑
Input Video
–
–
–
2.81
3.97
–
–
2.19
3.95
19.79
LucyEdit + HY-XL
cascade
0.16
0.12
0.12
2.88
3.76
0.19
0.25
2.79
2.27
9.81
LucyEdit + HY-XXL
cascade
0.16
0.12
0.12
2.88
3.76
0.20
0.29
2.89
2.15
13.20
VACE + Coherent
cascade
0.19
0.12
0.14
3.16
3.44
0.26
0.34
3.12
2.49
13.22
Table 2: CrossEditBench AV zero-shot scene editing results. Gray columns are the modality-specific FES subscores. Improvements due to our proposed techniques are underlined .
Overall
Visual
Audio
subscores
subscores
AV-FES ↑
V-FES ↑
Gain
Ins.
Cons.
A-FES ↑
Gain
Ins.
Cons.
Input
–
–
–
2.75 ± 0.24
4.71 ± 0.11
–
–
2.29 ± 0.25
4.68 ± 0.11
LucyEdit + HY-XL
0.11
0.14
0.12
2.61 ± 0.25
3.93 ± 0.20
0.08
0.16
2.17 ± 0.20
1.74 ± 0.20
LucyEdit + HY-XXL
0.14
0.17
0.15
2.74 ± 0.21
3.58 ± 0.20
0.12
0.17
2.22 ± 0.20
1.95 ± 0.21
CrossEdit-Base
0.26
0.25
0.34
2.85 ± 0.20
2.37 ± 0.19
0.27
0.32
2.89 ± 0.21
2.80 ± 0.21
Table 3: Ablation on zero-shot AV scene editing with human ratings (344 per system; Ins. and Cons. with 95% CIs). Improvements over CrossEdit-Base are underlined .
Figure 5: Human-rated AV-FES on CrossEditBench, by instruction category.
Speech
Video
Sync
Human Ratings
Audio
Video
WER ↓
SPK ↑
Time ↑
Input ↑
LSE-D ↓
LSE-C ↑
Ins. ↑
Cons. ↑
Ins. ↑
Cons. ↑
Input
-
-
-
-
-
-
3.29
4.39
3.29
4.50
F5-TTS
7.5
89.3
-
-
-
-
-
-
-
-
+ HuMo
-
-
99.1
67.4
9.12
5.88
3.30
2.97
2.60
2.10
+ StableAvatar
-
-
97.5
42.0
12.45
2.13
3.21
2.88
2.17
1.78
Table 4: Lip-synced speech editing results. Improvements due to our techniques are underlined .
VLM Judge
Quality
Text Align.
Temporal Cons.
Editing Quality ↑
Pick Score ↑
Frame ↑
Video ↑
CLIP ↑
DINO ↑
InsV2V
5.21
19.39
24.99
22.54
97.15
96.57
Lucy Edit
5.89
19.67
26.00
23.11
98.49
98.38
EditVerse
7.65
20.07
26.73
23.93
98.56
98.40
CrossEdit-Base
6.86
19.42
23.10
20.10
98.00
96.90
CrossEdit
6.85
19.82
25.21
22.26
98.43
97.52
Table 5: Video editing on EditVerseBench. Improvements from our techniques are underlined .
Param.
Add
Adjust
Extract
Replace
Remove
Background
Style
Hybrid
Action
Overall
BAGEL
14B
3.56
3.31
1.70
3.30
2.62
3.24
4.49
2.38
4.17
3.20
UniWorld-V1
12B
3.82
3.64
2.27
3.47
3.24
2.99
4.21
2.96
2.74
3.26
OmniGen2
7B
3.57
3.06
1.77
3.74
3.20
3.57
4.81
2.52
4.68
3.44
Qwen-Image
20B
4.38
4.16
3.43
4.66
4.14
4.38
4.81
3.82
4.69
4.27
CrossEdit-Base
12B
4.37
4.06
1.48
3.61
3.43
3.88
4.54
3.22
4.48
3.69
CrossEdit
12B
4.20
4.29
1.74
3.61
4.32
2.82
4.44
3.57
4.96
3.72
Table 6: Image editing results on ImgEdit (1–5, higher is better) . Dedicated image editing models are shown for reference, with scores from Wu et al. (2025a)
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Tokens
Batch Samples
Audio
Synthesized Audio Editing
0.5B
64
Simulated Audio Editing
2.5B
208
Text-to-Audio
0.3B
48
Speech Editing
1.5B
96
Video
Appendix
Table 7: SFT data mixture by total tokens and samples per batch. Video-to-video and lip-synced speech editing samples are replaced by AV masked reconstruction (Section 3.4 ) with 50% probability.
Figure 6: Example AV editing instructions, inputs, and outputs for CrossEditBench/CrossEdit.
Figure 7: Example AV editing instructions, inputs, and outputs for CrossEditBench/CrossEdit.
Figure 8: Instructions given to human annotators and the LLM Judges
Audio
Video
Ins.
Cons.
Ins.
Cons.
Input
3.29 ± 0.25
4.39 ± 0.18
3.29 ± 0.25
4.50 ± 0.16
F5-TTS+HuMo
3.30 ± 0.23
2.97 ± 0.22
2.60 ± 0.22
2.10 ± 0.18
F5-TTS+StableAvatar
3.21 ± 0.23
2.88 ± 0.22
2.17 ± 0.22
1.78 ± 0.20
CrossEdit-Base
3.98 ± 0.19
3.81 ± 0.19
4.05 ± 0.18
4.22 ± 0.17
+ Audio Editing
3.86 ± 0.20
3.82 ± 0.20
3.99 ± 0.19
4.35 ± 0.16
Appendix
Table 8: Full subjective results on IMDB lip-synced speech editing (1–5, mean ± 95% confidence interval).
Judge axes (1–5)
Faithful Edit Score
Vis. Ins.
Vis. Cons.
Aud. Ins.
Aud. Cons.
V-FES ↑
A-FES ↑
AV-FES ↑
Input
2.81 ± 0.19
3.97 ± 0.03
2.19 ± 0.10
3.95 ± 0.05
–
–
–
LucyEdit + HY-XL
2.88 ± 0.15
3.76 ± 0.16
2.79 ± 0.18
2.27 ± 0.07
0.12 ± 0.02
0.19 ± 0.04
0.16 ± 0.03
LucyEdit + HY-XXL
2.88 ± 0.16
3.76 ± 0.16
2.89 ± 0.17
2.15 ± 0.10
0.12 ± 0.04
0.20 ± 0.05
0.16 ± 0.03
VACE + Coherent
3.16 ± 0.17
3.44 ± 0.16
3.12 ± 0.17
2.49 ± 0.14
0.12 ± 0.04
0.26 ± 0.04
0.19 ± 0.03
AvED
2.81 ± 0.16
3.20 ± 0.13
2.49 ± 0.17
3.93 ± 0.05
0.09 ± 0.04
0.14 ± 0.05
0.12 ± 0.03
Appendix
Table 9: CrossEditBench results with Gemini 3.1 Pro as the judge, with 95% confidence intervals (mean ± half-width over the 100 clips). Means match Table 2 .
Judge axes (1–5)
Faithful Edit Score
Vis. Ins.
Vis. Cons.
Aud. Ins.
Aud. Cons.
V-FES ↑
A-FES ↑
AV-FES ↑
Input
2.69 ± 0.39
5.00 ± 0.00
3.32 ± 0.40
4.40 ± 0.28
–
–
–
LucyEdit + HY-XL
2.68 ± 0.38
4.10 ± 0.29
3.44 ± 0.38
2.47 ± 0.36
0.16 ± 0.09
0.19 ± 0.12
0.17 ± 0.07
LucyEdit + HY-XXL
3.88 ± 0.35
4.18 ± 0.25
3.84 ± 0.34
2.55 ± 0.38
0.50 ± 0.11
0.28 ± 0.13
0.41 ± 0.09
VACE + Coherent
3.85 ± 0.34
4.20 ± 0.24
3.73 ± 0.38
3.03 ± 0.38
0.33 ± 0.10
0.15 ± 0.10
0.25 ± 0.07
AvED
3.70 ± 0.37
4.48 ± 0.18
3.68 ± 0.36
4.23 ± 0.30
0.46 ± 0.12
0.26 ± 0.13
0.37 ± 0.09
Appendix
Table 10: CrossEditBench results with Qwen3-Omni as the judge, with 95% confidence intervals (mean ± half-width over the 100 clips). Means match Table 11 .
Judge axes (1–5)
Vis. Ins.
Vis. Cons.
Aud. Ins.
Aud. Cons.
V-FES ↑
A-FES ↑
AV-FES ↑
AV-FES ∣ edit ↑
No-gain ↓
Input
2.69
5.00
3.32
4.40
–
–
–
–
–
LucyEdit + HY-XL
2.68
4.10
3.44
2.47
0.16
0.19
0.17
0.17
75
LucyEdit + HY-XXL
3.88
4.18
3.84
2.55
0.50
0.28
0.41
0.38
55
VACE + Coherent
3.85
4.20
3.73
3.03
0.33
0.15
0.25
0.23
60
AvED
3.70
4.48
3.68
4.23
0.46
0.26
0.37
0.35
62
Appendix
Table 11: CrossEditBench results with Qwen3-Omni as the judge. Same rubric and protocol as Table 2 . AV-FES ∣ edit averages FES m only over clips whose instruction targets modality m . No-gain: number of clips (out of 100) with zero instruction gain.
All systems ( n=13 )
w/o input ( n=12 )
Score
Pearson
Spearman
Mean ∣ diff ∣
Gemini range
Qwen range
Pearson / Spearman
V-FES
0.96
0.95
0.356
0.00–0.29
0.00–0.80
0.95 / 0.93
A-FES
0.74
0.70
0.185
0.00–0.38
0.00–0.74
0.60 / 0.62
AV-FES
0.90
0.90
0.270
0.00–0.34
0.00–0.77
0.85 / 0.87
AV-FES ∣ edit
0.89
0.87
0.215
0.00–0.45
0.00–0.75
0.83 / 0.83
Appendix
Table 12: System-level agreement between Gemini 3.1 Pro and Qwen3-Omni judges. Computed over 13 systems, including development checkpoints not reported in the main paper and the unedited input, which scores 0 under both judges by construction. The last column excludes the input (12 systems). AV-FES ∣ edit averages FES m only over clips whose instruction targets modality m . Mean ∣ diff ∣ is in score units.
Score
Gemini 3.1 Pro
Qwen3-Omni
Vis. Ins.
0.93 / 1.00
0.82 / 1.00
Vis. Cons.
0.99 / 1.00
0.71 / 0.60
Aud. Ins.
0.71 / 0.50
0.94 / 0.97
Aud. Cons.
1.00 / 0.90
0.93 / 1.00
V-FES
0.92 / 1.00
0.90 / 1.00
A-FES
0.84 / 0.50
0.99 / 1.00
Appendix
Table 13: System-level agreement between LLM judges and human ratings (Pearson / Spearman) over the systems shared with Table 3 : n=4 for Gemini 3.1 Pro and n=5 for Qwen3-Omni.
Figure 9: Synthetic audio editing data reinforces visual consistency. Mined AV pairs are not perfectly consistent (here, the target has a text overlay and different lighting), causing the model to hallucinate unwanted changes (left). Our method corrects this (right).
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce AVE-Compass, a comprehensive benchmark with 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preserving, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Extensive evaluation shows that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. We further propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.
Yuqing Wen, Yukai Huang, Qianqian Xie +6
National University of Singapore · Beijing University of Posts and Telecommunications · NJU-LINK Team, Nanjing University +3
Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.
Abhinav Sharma, Sai Karthik Navuluru, Wang Wei +21
University of Massachusetts Amherst · University of Texas at Dallas · Virginia Tech +10
Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, we propose InstructAV2AV, the first end-to-end framework for instruction-guided audio-video joint editing. We first develop a scalable data synthesis pipeline and construct InsAVE-80K, the first large-scale audio-video editing dataset with high-quality source-to-target pairs. With this data foundation, we adapt an audio-video generation backbone to leverage its robust priors. We concatenate the audio-video input with noisy latent codes to anchor the source context, propose the source-instruction gated attention to improve instruction following and content preservation, and introduce a two-stage training strategy to effectively transfer these pre-trained priors. Extensive experiments demonstrate that InstructAV2AV outperforms state-of-the-art methods across 11 metrics spanning three aspects on two evaluation sets, highlighting its potential for controllable content creation. Project page: https://hjzheng.net/projects/InstructAV2AV/.
Haojie Zheng, Yixin Yang, Siqi Yang +2
Beijing Academy of Artificial Intelligence, Beijing, China · Peking University, Beijing, China