cs.SDSep 20, 2026

LiteCASS: A Lightweight End-to-End Network for Real-Time Stereo Cinematic Audio Source Separation

Authors: Yuanxin GuoQiang JiMengmei LiuYuhan LvNingning PanGongping Huang

Abstract

Cinematic audio source separation (CASS) decomposes a soundtrack into dialogue, music, and sound-effects (SFX) stems. Existing CASS methods, however, suffer from two critical limitations: they rely on heavily parameterized network architectures and GPU-class hardware, limiting their use in real-time and resource-constrained scenarios, and they are overwhelmingly designed for monaural signals, leaving the stereo scenario largely unexplored. We present LiteCASS, to our knowledge the first lightweight end-to-end network for real-time stereo CASS. LiteCASS combines deterministic STFT subband rearrangement with two jointly trained compact U-Nets: the first extracts dialogue, and the second separates music and SFX from the predicted non-speech component. A multi-task waveform-domain L1 loss supervises all stems. On a spatialized stereo extension of DnR v3, LiteCASS-K8 uses only 1.06M parameters and 0.72G MACs per second, while achieving the highest averaged SI-SDR among the compared CASS baselines.

Explore similar work

Sep 15, 2025cs.SD

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

Text-guided sound separation enables flexible audio editing, assistive listening, and open-domain source extraction, but systems such as AudioSep remain too expensive for low-latency edge or codec-mediated deployment. Existing neural audio codec separators are efficient, yet largely restricted to fixed stems or closed taxonomies. We introduce CodecSep, a prompt-driven universal sound separation framework that extracts sources directly in neural audio codec latent space. CodecSep combines a frozen DAC backbone with a lightweight FiLM-conditioned Transformer masker driven by CLAP text embeddings, enabling open-vocabulary separation while preserving codec-native efficiency. Across dnr-v2 and five open-domain benchmarks, CodecSep consistently improves over AudioSep in SI-SDR, remains competitive in ViSQOL, and achieves clear gains in human MOS-LQS. Controlled analyses show that fine-grained prompts outperform coarse labels, and that explicit latent masking is substantially more effective than decoder-style latent generation in codec space. Qualitative diagnostics show that neural audio codec latents retain source-dependent structure, which CodecSep exploits mainly through channel-wise source-conditioned modulation. CodecSep also provides a practical code-stream deployment path. When audio is transmitted as neural audio codec codes, CodecSep maps codes to embeddings, separates directly in codec space, and outputs waveforms or re-quantized codes, avoiding the decode-separate-re-encode loop. In this regime, CodecSep requires only 1.35 GMACs end-to-end: about 54 times less compute than AudioSep in the same pipeline and 25 times lower separator-only compute, with much lower latency and memory. More broadly, CodecSep offers a blueprint for codec-native downstream audio processing.
Adhiraj Banerjee, Vipul Arora
Sep 14, 2026cs.SD

Real-Time Music Source Separation on a Low-Power Audio DSP

Real-time music source separation is validated on desktop CPUs and GPUs. Does any published system fit the embedded audio hardware it targets? On a commercial audio DSP (2 MB SRAM, 2.07 GMAC/s measured), none does, and the constraints eliminate different models: memory rules out the 16-51 M parameter TasNet/X-UMX family, per-frame compute rules out RT-STT, needing 5.5x the available MAC rate. Parameter count predicts neither: weight reuse spans 1x to 345x. We then build one that fits. Training on continuous rather than block-padded convolution context proves essential: a model scoring 3.93 dB block-wise otherwise collapses to silence within 2 s frame-by-frame. A gated complex FIR deep filter adds a latency knob, gaining 0.38 dB even when strictly causal. It reaches 4.70 dB cSDR on MUSDB18-HQ and runs in 10.43 ms of an 11.6 ms hop, 0.5-0.7 dB behind systems that do not fit.
Jianan Li, Li Liu, Ken Malsky +1
Jul 25, 2026cs.SD

Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models

Music Source Separation (MSS), the task of recovering individual sound components (stems) from a polyphonic mixture, is central to applications ranging from karaoke and remixing to audio restoration and content production. The separation quality depends on engineering decisions across the entire pipeline: model choice, training data preparation and augmentation, loss function and metrics choice, training configuration, validation, and post-processing. This paper presents MSST (Music-Source-Separation-Training) - a universal open-source framework for MSS tasks, which unifies training, validation, and inference for a broad range of modern demixing model families under a single, configuration-driven interface. The framework supports various model architectures, data preprocessing and augmentations, multiple loss functions and evaluation metrics, which helps with fast iterations and ablation studies. Additionally, the framework supports a range of practical techniques that improve separation quality, such as sliding-window inference with cross-fading, test-time augmentation, model ensembling, and fine-tuning via Low-Rank Adaptation (LORA). Our ablation studies demonstrate improvements of MSS using the above techniques. By consolidating these components into a reproducible, YAML-configurable framework, MSST lowers the barrier to systematic experimentation and enables rapid iteration from idea to verifiable result.
Roman Solovyev, Ilya Kiselev, Alexander Stempkovskiy +1