AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes
Authors: Weihan Xu, Kan Jen Cheng, Koichi Saito, Jingyu Shi, Tingle Li, Yisi Liu, Liming Wang, Masato Ishii, +3 more
Organizations: University of Washington · Massachusetts Institute of Technology · University of Maryland, College Park · Sony Research · Independent Researcher · University of California, Berkeley
Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.
Figures & tables
Figure 1: Audiovisual event addition and removal with AVIO. (a) A donkey and its braying are added to the scene with an optional reference frame to specify initial placement and appearance (b) A lion and its roaring are removed while the vehicle and its sound are preserved.
Figure 2: Overview of the AVIOBench data construction pipeline. For each clip, we identify the primary sounding object O , remove it visually via segmentation and inpainting, and remove its sound via audio separation. Each stage applies automatic filters (yellow) to discard unreliable candidates, and the edited streams are jointly refined to form X∖O .
Figure 3: Overview of AVIO. We adapt a pretrained text-to-audiovisual DiT by conditioning it on source audiovisual latents. To support addition and removal within a unified model, we introduce a reference-guided curriculum that progressively increases the dropout probability of the first target frame during training.
Figure 4: Qualitative comparison of audiovisual object removal (left) and addition (right). Each row shows video frames and audio waveforms for the corresponding model (left). † CoherentAVEdit uses oracle target video; ‡ AVI-Edit uses oracle target audio. Red and Green highlight removal targets and added content, respectively. AVIO successfully edits objects and their sounds while preserving the surrounding scene and matching the appearance and placement of the reference object for addition.
Task
Input
Model
PSNR ↑
SSIM ↑
LPIPS ↓
FVD ↓
Removal
Text
LGVI ( 2024 )
16.22
0.55
0.40
402.69
VideoPainter ( 2025 )
18.43
0.68
0.30
362.35
ROSE ( 2025 )
20.30
0.72
0.24
176.05
InstructAV2AV ( 2026b )
18.88
0.68
0.29
180.10
AVI-Edit ( 2026a )
16.65
0.64
0.32
580.48
Ours
20.06
0.66
0.28
190.50
Table 1: Visual, audio, and audiovisual editing evaluation. Reference-based metrics use pseudo-ground-truth targets. VES, AES, and AVC are Gemini 3.1 Pro Yes rates (%). Best and second-best results per task are bold and underlined, using unrounded scores and excluding Oracle / Oracle.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Visual quality
Audio quality
Audiovisual evaluation
Task
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
FVD ↓
FAD ↓
IS ↑
KL ↓
LSD ↓
DeSync ↓
VES ↑
AES ↑
AVC ↑
Removal
AVIO-XAttn
14.51
0.419
0.541
525
2.383
4.43
1.960
15.01
0.913
62.5
20.0
90.0
AVIO-NoRef
20.09
0.665
0.278
206
1.582
6.58
1.373
13.03
0.846
40.5
13.5
84.0
AVIO
20.06
0.661
0.279
191
1.972
6.48
1.355
12.81
0.869
37.5
9.5
88.5
Addition
AVIO-XAttn
12.68
0.328
0.529
977
8.210
4.98
2.216
15.02
1.054
10.0
71.0
85.5
AVIO-NoRef
16.65
0.568
0.319
538
7.145
6.46
1.451
12.32
0.993
16.5
77.0
81.0
Appendix
Table 2: Ablation study of source conditioning and reference guidance. Each variant is trained until validation performance no longer improves, with checkpoint selection based on the validation set. The w/o reference variant omits reference frames during training and inference. Source attention replaces feature modulation with attention to the source. For addition, Ours and Source attention receive an edited reference frame. Removal uses no reference frame. Due to resource constraints, Gemini-based evaluation (VES, AES, and AVC) is restricted to a subset of 200 test clips. Other metrics use the full test set. Best and second-best results per task are bold and underlined.
Annotations
Dataset
Duration (h)
Modalities
Event focus
Instruction
Mask
Source
Target
OpenVE-3M ( 2026 )
–
V
–
✓
–
✓
✓
Goku ( 2026b )
–
V
–
✓
– †
✓
✓
AUDIT ( 2023 )
–
A
–
✓
–
✓
✓
AvED-Bench ( 2025 )
0.31
A+V
✓
–
–
✓
–
Object-AVEdit ( 2025 )
–
A+V
✓
–
–
✓
–
Appendix
Table 3: Comparison of datasets by duration, modality, audio focus, and editing supervision. Event focus indicates a focus on non-speech sounding-object edits. ✓ : documented; –: not established or not applicable. A: audio; V: video. Duration counts each clip or editing pair once, without summing source and target durations; unverified durations are denoted by –.
Stage
In
Out
Kept
Cum.
Visual filtering
Visibility †
114,544
≥112,534
≥98.2%
≥98.2%
Inpainting
≥105,333
87,765
≤83.3%
76.6%
Removal check ‡
82,608
38,031
46.0%
33.2%
Audio filtering
Complete pair
38,031
35,669
93.8%
31.1%
Appendix
Table 4: Dataset retention across filtering stages. In and Out denote clip counts. Kept is stage-wise retention; Cum. is retention relative to the 114,544 initial clips. Bounds reflect incomplete intermediate records.
Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual inconsistency. Existing video inpainting models operate solely on pixels, while audio editing models, especially for the sound removal task, are typically driven by text and therefore rely on limited single-modal control, which is less effective than multimodal guidance that provides stronger semantic grounding and temporal synchronization cues. In this paper, we present Text-Visual Guided Sound Removal (TV-AudioRemover), a target sound removal framework that leverages the visually edited video together with a natural-language instruction to suppress the sound associated with the removed visual object from the original audio mixture. To acquire high-quality training data, we devise a pipeline to construct a million-scale dataset of single-object audio-visual aligned samples, from which we synthesize mixture-target pairs customized for model training. To effectively leverage visual context and follow instruction intent, we augment the model architecture with task tokens, generalizable instruction modeling, and modality-specific global guidance. We further adopt multi-task training to strengthen task-role comprehension, and employ a hard-mixture curriculum that leverages semantically similar acoustic mixtures during fine-tuning to enhance fine-grained source discrimination. To support evaluation, we present AV-Remove-Bench, a comprehensive audio-visual object removal benchmark, along with dedicated objective metrics and an MLLM-based evaluation protocol. Experiments demonstrate that our method achieves state-of-the-art performance on both subjective and objective metrics. Project page: https://yjx-research.github.io/TV-AudioRemover/.
Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, we propose InstructAV2AV, the first end-to-end framework for instruction-guided audio-video joint editing. We first develop a scalable data synthesis pipeline and construct InsAVE-80K, the first large-scale audio-video editing dataset with high-quality source-to-target pairs. With this data foundation, we adapt an audio-video generation backbone to leverage its robust priors. We concatenate the audio-video input with noisy latent codes to anchor the source context, propose the source-instruction gated attention to improve instruction following and content preservation, and introduce a two-stage training strategy to effectively transfer these pre-trained priors. Extensive experiments demonstrate that InstructAV2AV outperforms state-of-the-art methods across 11 metrics spanning three aspects on two evaluation sets, highlighting its potential for controllable content creation. Project page: https://hjzheng.net/projects/InstructAV2AV/.
Haojie Zheng, Yixin Yang, Siqi Yang +2
Beijing Academy of Artificial Intelligence, Beijing, China · Peking University, Beijing, China
While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets and benchmarks. To bridge this gap, we present JAVEdit-100k, the first large-scale, high-quality dataset tailored for instruction-guided joint audio-visual editing. Focusing on human-centric videos, JAVEdit-100k comprises approximately 100K editing triplets spanning five distinct categories, including subject editing and speech editing. This dataset is rigorously constructed via four meticulously designed generation pipelines, seamlessly paired with an agent-in-the-loop quality control mechanism. Furthermore, to address the lack of standardized evaluation within the field, we introduce JAVEditBench, a comprehensive benchmark featuring curated source videos and human-aligned instructions across all editing categories. Finally, we propose JAVEdit, a pioneering baseline model for instruction-guided joint audio-visual editing. Experiments show that \model\ outperforms all baselines on five of six evaluation metrics.
Yinan Chen, Chuming Lin, Zhennan Chen +12
1Zhejiang University · 2Tencent Youtu Lab · 3Nanjing University +3