cs.SDOct 7, 2026

A Study on Improving Multi-class Audio Source Separation Via Decoupled CLAP Query Optimization and an Automated Data Engine

Authors: Amirhossein Hajavi, Hanhee Lee, Pushya Jain, Sky Qiao, Emmanuel Ko, Yuanhao Yu, Irina Kezele

Organizations: Noah’s Ark Lab, Huawei Technologies Canada co.

Abstract

Language-queried audio source separation (LASS) enables extracting any sound source using natural language. However, adapting LASS models to application-specific sound classes is challenging due to noisy training data and limited semantic coverage of the CLAP-based control signals. We propose a framework comprised of an automated data engine for training-data curation and a two-stage optimization process for class-specific CLAP control signals. Our objective evaluations across seven sound classes show that data refinement and control signal optimization consistently improve source separation performance. Subjective evaluation with 17 participants further demonstrates perceptual improvements of the model trained with optimized control signals over baseline and similar commercial models.

Figures & tables

Explore similar work

CardsList
  1. Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

    Jul 25, 2026Roman Solovyev, Ilya Kiselev, Alexander Stempkovskiy +1Speech SeparationUnified Framework

  2. CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

    Sep 15, 2025Adhiraj Banerjee, Vipul AroraNeural Audio CodecsSpeech Separation

  3. Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models

    Jun 8, 2026Yuxuan Chen, Haoyuan Yu, Peize HeSpeech SeparationAudio Understanding