cs.SDOct 7, 2026

CrossEdit: Cross-Modal Training Enables Rich Audio-Visual Editing

Authors: William Chen, Prem Seetharaman, Ke Chen, Oriol Nieto, Kevin Duarte, Siddharth Srinivasan Iyer, Mamshad Nayeem Rizve, Zhiwen Cao, +5 more

Organizations: Adobe Research · Carnegie Mellon University

Abstract

Multimedia editing requires generative models to interpret complex, compositional instructions and modify only the desired elements across modalities, a capability largely beyond existing systems. Current approaches are typically limited to simple, single-attribute edits within one modality, since diverse, high-quality editing pairs are hard to obtain, especially for cross-modal tasks where aligned audiovisual (AV) data is scarce. We present CrossEdit, a unified omni-modal editing model for images, audio, and video that performs AV movie scene edits zero-shot through cross-modal transfer. Our key observation is that complex instruction-following, once learned in any modality, generalizes to others. We exploit the additive nature of acoustic signals to procedurally generate large-scale audio editing pairs with compositional instructions, and show that fine-tuning on this synthetic data and self-supervised AV masked reconstruction, alongside a targeted set of curated cross-modal tasks, induces instruction-following that transfers zero-shot to unseen modality and instruction combinations. To evaluate this capability, we release CrossEditBench, a human-annotated benchmark of AV edits on movie scenes, and propose AV-FES, a metric that jointly scores instruction following and consistency. We show that our proposed techniques improve performance on zero-shot AV movie scene editing, while maintaining or improving performance on lip-synced speech editing and standard image, video, and audio editing benchmarks. Demos are available at https://wanchichen.github.io/crossedit/.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

    Jul 17, 2026Yuqing Wen, Yukai Huang, Qianqian Xie +6Video EditingAudio-Visual Consistency

  2. Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy

    Sep 28, 2026Abhinav Sharma, Sai Karthik Navuluru, Wang Wei +21Audio-Video GenerationAudio-Visual Consistency

  3. InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

    May 18, 2026Haojie Zheng, Yixin Yang, Siqi Yang +2Audio-Video GenerationVideo Editing