cs.AISep 29, 2026

AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes

Authors: Weihan Xu, Kan Jen Cheng, Koichi Saito, Jingyu Shi, Tingle Li, Yisi Liu, Liming Wang, Masato Ishii, +3 more

Organizations: University of Washington · Massachusetts Institute of Technology · University of Maryland, College Park · Sony Research · Independent Researcher · University of California, Berkeley

Abstract

Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

    May 18, 2026Haojie Zheng, Yixin Yang, Siqi Yang +2Audio-Video GenerationVideo Editing

  2. JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

    Jun 2, 2026Yinan Chen, Chuming Lin, Zhennan Chen +12Video EditingVideo Dataset