cs.SDAug 31, 2026

Decoupled Latent Flow Matching for Few-Step Joint Vocal-Accompaniment Separation

Authors: Lishi ZuoYouzhi TuLu YiZezhong JinChongxin GanMan-Wai MakKongAik Lee

Organizations: Dept. of Electrical and Electronic Engineering, Hong Kong Polytechnic University, Hong Kong SAR, China

Abstract

Generative modeling provides a flexible way to model mixture-conditioned source distributions, but iterative diffusion and flow matching models are costly for long music signals. This paper studies joint vocal-accompaniment separation through latent flow matching, where a pretrained variational autoencoder (VAE) maps mixtures and sources into a compact latent space and a flow matching model generates vocal and accompaniment latents jointly. The proposed framework decouples semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder. To reduce sampling cost, we further apply latent adversarial post-training inspired by Flow2GAN for few-step generation. Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget.

Explore similar work

CardsList
  1. SURF: Separation via Unsupervised Remixing Flow

    Jun 3, 2026Henry Li, Robin Scheibler, Efthymios Tzinis +3Speech SeparationMix