CrossEdit: Cross-Modal Training Enables Rich Audio-Visual Editing
Organizations: Adobe Research · Carnegie Mellon University
Abstract
Multimedia editing requires generative models to interpret complex, compositional instructions and modify only the desired elements across modalities, a capability largely beyond existing systems. Current approaches are typically limited to simple, single-attribute edits within one modality, since diverse, high-quality editing pairs are hard to obtain, especially for cross-modal tasks where aligned audiovisual (AV) data is scarce. We present CrossEdit, a unified omni-modal editing model for images, audio, and video that performs AV movie scene edits zero-shot through cross-modal transfer. Our key observation is that complex instruction-following, once learned in any modality, generalizes to others. We exploit the additive nature of acoustic signals to procedurally generate large-scale audio editing pairs with compositional instructions, and show that fine-tuning on this synthetic data and self-supervised AV masked reconstruction, alongside a targeted set of curated cross-modal tasks, induces instruction-following that transfers zero-shot to unseen modality and instruction combinations. To evaluate this capability, we release CrossEditBench, a human-annotated benchmark of AV edits on movie scenes, and propose AV-FES, a metric that jointly scores instruction following and consistency. We show that our proposed techniques improve performance on zero-shot AV movie scene editing, while maintaining or improving performance on lip-synced speech editing and standard image, video, and audio editing benchmarks. Demos are available at https://wanchichen.github.io/crossedit/.
Figures & tables
| Task | Example Instruction | Pre-train | SFT | Eval. |
| Editing | ||||
| General image editing (I2I) | “Make her hair blue” | ✓ | ✓ | § 5.3 |
| General video editing (V2V) | “Turn the car red” | ✓ | ✓ | § 5.2 |
| Speech editing † (A2A) | “Change the speech to ‘See you soon”’ | ✓ | ✓ | – |
| Complex audio editing (A2A) | “Remove the music; make the rain louder” | ✗ | ✓ | § D |
| Lip-synced speech editing (AV2AV) | “Change the speech to ‘We made it out”’ | ✓ | ✓ | § 5.1 |
| Overall | Visual | Audio | Alignment | ||||||||
| subscores | subscores | ||||||||||
| Type | AV-FES | V-FES | Gain | Ins. | Cons. | A-FES | Gain | Ins. | Cons. | ImageBind | |
| Input Video | – | – | – | 2.81 | 3.97 | – | – | 2.19 | 3.95 | 19.79 | |
| LucyEdit + HY-XL | cascade | 0.16 | 0.12 | 0.12 | 2.88 | 3.76 | 0.19 | 0.25 | 2.79 | 2.27 | 9.81 |
| LucyEdit + HY-XXL | cascade | 0.16 | 0.12 | 0.12 | 2.88 | 3.76 | 0.20 | 0.29 | 2.89 | 2.15 | 13.20 |
| VACE + Coherent | cascade | 0.19 | 0.12 | 0.14 | 3.16 | 3.44 | 0.26 | 0.34 | 3.12 | 2.49 | 13.22 |
| Overall | Visual | Audio | |||||||
|---|---|---|---|---|---|---|---|---|---|
| subscores | subscores | ||||||||
| AV-FES | V-FES | Gain | Ins. | Cons. | A-FES | Gain | Ins. | Cons. | |
| Input | – | – | – | 2.75 0.24 | 4.71 0.11 | – | – | 2.29 0.25 | 4.68 0.11 |
| LucyEdit + HY-XL | 0.11 | 0.14 | 0.12 | 2.61 0.25 | 3.93 0.20 | 0.08 | 0.16 | 2.17 0.20 | 1.74 0.20 |
| LucyEdit + HY-XXL | 0.14 | 0.17 | 0.15 | 2.74 0.21 | 3.58 0.20 | 0.12 | 0.17 | 2.22 0.20 | 1.95 0.21 |
| CrossEdit-Base | 0.26 | 0.25 | 0.34 | 2.85 0.20 | 2.37 0.19 | 0.27 | 0.32 | 2.89 0.21 | 2.80 0.21 |
| Speech | Video | Sync | Human Ratings | |||||||
| Audio | Video | |||||||||
| WER | SPK | Time | Input | LSE-D | LSE-C | Ins. | Cons. | Ins. | Cons. | |
| Input | - | - | - | - | - | - | 3.29 | 4.39 | 3.29 | 4.50 |
| F5-TTS | 7.5 | 89.3 | - | - | - | - | - | - | - | - |
| + HuMo | - | - | 99.1 | 67.4 | 9.12 | 5.88 | 3.30 | 2.97 | 2.60 | 2.10 |
| + StableAvatar | - | - | 97.5 | 42.0 | 12.45 | 2.13 | 3.21 | 2.88 | 2.17 | 1.78 |
| VLM Judge | Quality | Text Align. | Temporal Cons. | |||
| Editing Quality | Pick Score | Frame | Video | CLIP | DINO | |
| InsV2V | 5.21 | 19.39 | 24.99 | 22.54 | 97.15 | 96.57 |
| Lucy Edit | 5.89 | 19.67 | 26.00 | 23.11 | 98.49 | 98.38 |
| EditVerse | 7.65 | 20.07 | 26.73 | 23.93 | 98.56 | 98.40 |
| CrossEdit-Base | 6.86 | 19.42 | 23.10 | 20.10 | 98.00 | 96.90 |
| CrossEdit | 6.85 | 19.82 | 25.21 | 22.26 | 98.43 | 97.52 |
| Param. | Add | Adjust | Extract | Replace | Remove | Background | Style | Hybrid | Action | Overall | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| BAGEL | 14B | 3.56 | 3.31 | 1.70 | 3.30 | 2.62 | 3.24 | 4.49 | 2.38 | 4.17 | 3.20 |
| UniWorld-V1 | 12B | 3.82 | 3.64 | 2.27 | 3.47 | 3.24 | 2.99 | 4.21 | 2.96 | 2.74 | 3.26 |
| OmniGen2 | 7B | 3.57 | 3.06 | 1.77 | 3.74 | 3.20 | 3.57 | 4.81 | 2.52 | 4.68 | 3.44 |
| Qwen-Image | 20B | 4.38 | 4.16 | 3.43 | 4.66 | 4.14 | 4.38 | 4.81 | 3.82 | 4.69 | 4.27 |
| CrossEdit-Base | 12B | 4.37 | 4.06 | 1.48 | 3.61 | 3.43 | 3.88 | 4.54 | 3.22 | 4.48 | 3.69 |
| CrossEdit | 12B | 4.20 | 4.29 | 1.74 | 3.61 | 4.32 | 2.82 | 4.44 | 3.57 | 4.96 | 3.72 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Tokens | Batch Samples |
|---|---|---|
| Audio | ||
| Synthesized Audio Editing | 0.5B | 64 |
| Simulated Audio Editing | 2.5B | 208 |
| Text-to-Audio | 0.3B | 48 |
| Speech Editing | 1.5B | 96 |
| Video |
| Audio | Video | |||
| Ins. | Cons. | Ins. | Cons. | |
| Input | 3.29 0.25 | 4.39 0.18 | 3.29 0.25 | 4.50 0.16 |
| F5-TTS+HuMo | 3.30 0.23 | 2.97 0.22 | 2.60 0.22 | 2.10 0.18 |
| F5-TTS+StableAvatar | 3.21 0.23 | 2.88 0.22 | 2.17 0.22 | 1.78 0.20 |
| CrossEdit-Base | 3.98 0.19 | 3.81 0.19 | 4.05 0.18 | 4.22 0.17 |
| + Audio Editing | 3.86 0.20 | 3.82 0.20 | 3.99 0.19 | 4.35 0.16 |
| Judge axes (1–5) | Faithful Edit Score | ||||||
| Vis. Ins. | Vis. Cons. | Aud. Ins. | Aud. Cons. | V-FES | A-FES | AV-FES | |
| Input | 2.81 0.19 | 3.97 0.03 | 2.19 0.10 | 3.95 0.05 | – | – | – |
| LucyEdit + HY-XL | 2.88 0.15 | 3.76 0.16 | 2.79 0.18 | 2.27 0.07 | 0.12 0.02 | 0.19 0.04 | 0.16 0.03 |
| LucyEdit + HY-XXL | 2.88 0.16 | 3.76 0.16 | 2.89 0.17 | 2.15 0.10 | 0.12 0.04 | 0.20 0.05 | 0.16 0.03 |
| VACE + Coherent | 3.16 0.17 | 3.44 0.16 | 3.12 0.17 | 2.49 0.14 | 0.12 0.04 | 0.26 0.04 | 0.19 0.03 |
| AvED | 2.81 0.16 | 3.20 0.13 | 2.49 0.17 | 3.93 0.05 | 0.09 0.04 | 0.14 0.05 | 0.12 0.03 |
| Judge axes (1–5) | Faithful Edit Score | ||||||
| Vis. Ins. | Vis. Cons. | Aud. Ins. | Aud. Cons. | V-FES | A-FES | AV-FES | |
| Input | 2.69 0.39 | 5.00 0.00 | 3.32 0.40 | 4.40 0.28 | – | – | – |
| LucyEdit + HY-XL | 2.68 0.38 | 4.10 0.29 | 3.44 0.38 | 2.47 0.36 | 0.16 0.09 | 0.19 0.12 | 0.17 0.07 |
| LucyEdit + HY-XXL | 3.88 0.35 | 4.18 0.25 | 3.84 0.34 | 2.55 0.38 | 0.50 0.11 | 0.28 0.13 | 0.41 0.09 |
| VACE + Coherent | 3.85 0.34 | 4.20 0.24 | 3.73 0.38 | 3.03 0.38 | 0.33 0.10 | 0.15 0.10 | 0.25 0.07 |
| AvED | 3.70 0.37 | 4.48 0.18 | 3.68 0.36 | 4.23 0.30 | 0.46 0.12 | 0.26 0.13 | 0.37 0.09 |
| Judge axes (1–5) | |||||||||
| Vis. Ins. | Vis. Cons. | Aud. Ins. | Aud. Cons. | V-FES | A-FES | AV-FES | AV-FES edit | No-gain | |
| Input | 2.69 | 5.00 | 3.32 | 4.40 | – | – | – | – | – |
| LucyEdit + HY-XL | 2.68 | 4.10 | 3.44 | 2.47 | 0.16 | 0.19 | 0.17 | 0.17 | 75 |
| LucyEdit + HY-XXL | 3.88 | 4.18 | 3.84 | 2.55 | 0.50 | 0.28 | 0.41 | 0.38 | 55 |
| VACE + Coherent | 3.85 | 4.20 | 3.73 | 3.03 | 0.33 | 0.15 | 0.25 | 0.23 | 60 |
| AvED | 3.70 | 4.48 | 3.68 | 4.23 | 0.46 | 0.26 | 0.37 | 0.35 | 62 |
| All systems ( ) | w/o input ( ) | |||||
|---|---|---|---|---|---|---|
| Score | Pearson | Spearman | Mean diff | Gemini range | Qwen range | Pearson / Spearman |
| V-FES | 0.96 | 0.95 | 0.356 | 0.00–0.29 | 0.00–0.80 | 0.95 / 0.93 |
| A-FES | 0.74 | 0.70 | 0.185 | 0.00–0.38 | 0.00–0.74 | 0.60 / 0.62 |
| AV-FES | 0.90 | 0.90 | 0.270 | 0.00–0.34 | 0.00–0.77 | 0.85 / 0.87 |
| AV-FES edit | 0.89 | 0.87 | 0.215 | 0.00–0.45 | 0.00–0.75 | 0.83 / 0.83 |
| Score | Gemini 3.1 Pro | Qwen3-Omni |
|---|---|---|
| Vis. Ins. | 0.93 / 1.00 | 0.82 / 1.00 |
| Vis. Cons. | 0.99 / 1.00 | 0.71 / 0.60 |
| Aud. Ins. | 0.71 / 0.50 | 0.94 / 0.97 |
| Aud. Cons. | 1.00 / 0.90 | 0.93 / 1.00 |
| V-FES | 0.92 / 1.00 | 0.90 / 1.00 |
| A-FES | 0.84 / 0.50 | 0.99 / 1.00 |
| StoryGenEval | AudioCapsEdit | Movie Separation | ||||
| FLAM | editFLAM | editFLAM | FLAM | editFLAM | SAM-J | |
| SAM-Audio | - | - | - | 18.97 | 4.12 | 2.96 |
| AudioChat | 11.7 | 18.6 | 15.5 | - | - | - |
| CrossEdit | 6.64 | 23.1 | 40.3 | 7.08 | 4.29 | 3.41 |