cs.SDSep 24, 2026

Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing

Authors: Ilpo Viertola, Giulio Cengarle, Gouthaman KV, Daniel Arteaga, Lie Lu

Organizations: Dolby Laboratories · Tampere University

Abstract

We introduce Spot, Separate, and Enhance (SSE), the first multimodal, user-guided generative model for audio remixing and enhancement. SSE enhances video content by rebalancing the audio, removing unwanted audio sources, and reducing reverberation, guided by both video and textual descriptions. To support its training and evaluation, we propose DegradedMix, a new dataset built on the audio remixing benchmark MuddyMix. We also adopt evaluation metrics from generative modeling, which better capture the creative nature of remixing than standard reconstruction-based metrics. SSE outperforms existing baselines in both controllability and remixing quality, as shown by extensive experiments. Project page: https://sse-ai.notion.site

Figures & tables

Explore similar work

CardsList
  1. Soundwich: Video Generation with Layered and Controllable Audio

    Sep 30, 2026Zhuo Ning, AmirHossein Naghi Razlighi, Sagi Polaczek +2Audio-Video GenerationVideo Generation

  2. Diff2Mix: Controllable Music Mixing via Diffusion Models and Differentiable Audio Effects

    Aug 5, 2026Yisu Zong, Jinjie Shi, Joshua ReissAudio EditingMix

  3. Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text

    May 27, 2026Jiahao Mei, Heinrich Dinkel, Yadong Niu +7Modern Generative Audio ModelsAcoustic Latent Space