AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes
Authors: Weihan Xu, Kan Jen Cheng, Koichi Saito, Jingyu Shi, Tingle Li, Yisi Liu, Liming Wang, Masato Ishii, +3 more
Organizations: University of Washington · Massachusetts Institute of Technology · University of Maryland, College Park · Sony Research · Independent Researcher · University of California, Berkeley
Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.
Figures & tables
Figure 1: Audiovisual event addition and removal with AVIO. (a) A donkey and its braying are added to the scene with an optional reference frame to specify initial placement and appearance (b) A lion and its roaring are removed while the vehicle and its sound are preserved.
Figure 2: Overview of the AVIOBench data construction pipeline. For each clip, we identify the primary sounding object O , remove it visually via segmentation and inpainting, and remove its sound via audio separation. Each stage applies automatic filters (yellow) to discard unreliable candidates, and the edited streams are jointly refined to form X∖O .
Figure 3: Overview of AVIO. We adapt a pretrained text-to-audiovisual DiT by conditioning it on source audiovisual latents. To support addition and removal within a unified model, we introduce a reference-guided curriculum that progressively increases the dropout probability of the first target frame during training.
Figure 4: Qualitative comparison of audiovisual object removal (left) and addition (right). Each row shows video frames and audio waveforms for the corresponding model (left). † CoherentAVEdit uses oracle target video; ‡ AVI-Edit uses oracle target audio. Red and Green highlight removal targets and added content, respectively. AVIO successfully edits objects and their sounds while preserving the surrounding scene and matching the appearance and placement of the reference object for addition.
Task
Input
Model
PSNR ↑
SSIM ↑
LPIPS ↓
FVD ↓
Removal
Text
LGVI ( 2024 )
16.22
0.55
0.40
402.69
VideoPainter ( 2025 )
18.43
0.68
0.30
362.35
ROSE ( 2025 )
20.30
0.72
0.24
176.05
InstructAV2AV ( 2026b )
18.88
0.68
0.29
180.10
AVI-Edit ( 2026a )
16.65
0.64
0.32
580.48
Ours
20.06
0.66
0.28
190.50
Table 1: Visual, audio, and audiovisual editing evaluation. Reference-based metrics use pseudo-ground-truth targets. VES, AES, and AVC are Gemini 3.1 Pro Yes rates (%). Best and second-best results per task are bold and underlined, using unrounded scores and excluding Oracle / Oracle.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Visual quality
Audio quality
Audiovisual evaluation
Task
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
FVD ↓
FAD ↓
IS ↑
KL ↓
LSD ↓
DeSync ↓
VES ↑
AES ↑
AVC ↑
Removal
AVIO-XAttn
14.51
0.419
0.541
525
2.383
4.43
1.960
15.01
0.913
62.5
20.0
90.0
AVIO-NoRef
20.09
0.665
0.278
206
1.582
6.58
1.373
13.03
0.846
40.5
13.5
84.0
AVIO
20.06
0.661
0.279
191
1.972
6.48
1.355
12.81
0.869
37.5
9.5
88.5
Addition
AVIO-XAttn
12.68
0.328
0.529
977
8.210
4.98
2.216
15.02
1.054
10.0
71.0
85.5
AVIO-NoRef
16.65
0.568
0.319
538
7.145
6.46
1.451
12.32
0.993
16.5
77.0
81.0
Appendix
Table 2: Ablation study of source conditioning and reference guidance. Each variant is trained until validation performance no longer improves, with checkpoint selection based on the validation set. The w/o reference variant omits reference frames during training and inference. Source attention replaces feature modulation with attention to the source. For addition, Ours and Source attention receive an edited reference frame. Removal uses no reference frame. Due to resource constraints, Gemini-based evaluation (VES, AES, and AVC) is restricted to a subset of 200 test clips. Other metrics use the full test set. Best and second-best results per task are bold and underlined.
Annotations
Dataset
Duration (h)
Modalities
Event focus
Instruction
Mask
Source
Target
OpenVE-3M ( 2026 )
–
V
–
✓
–
✓
✓
Goku ( 2026b )
–
V
–
✓
– †
✓
✓
AUDIT ( 2023 )
–
A
–
✓
–
✓
✓
AvED-Bench ( 2025 )
0.31
A+V
✓
–
–
✓
–
Object-AVEdit ( 2025 )
–
A+V
✓
–
–
✓
–
Appendix
Table 3: Comparison of datasets by duration, modality, audio focus, and editing supervision. Event focus indicates a focus on non-speech sounding-object edits. ✓ : documented; –: not established or not applicable. A: audio; V: video. Duration counts each clip or editing pair once, without summing source and target durations; unverified durations are denoted by –.
Stage
In
Out
Kept
Cum.
Visual filtering
Visibility †
114,544
≥112,534
≥98.2%
≥98.2%
Inpainting
≥105,333
87,765
≤83.3%
76.6%
Removal check ‡
82,608
38,031
46.0%
33.2%
Audio filtering
Complete pair
38,031
35,669
93.8%
31.1%
Appendix
Table 4: Dataset retention across filtering stages. In and Out denote clip counts. Kept is stage-wise retention; Cum. is retention relative to the 114,544 initial clips. Bounds reflect incomplete intermediate records.