Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training
Authors: Houmin Sun, Zi Hu, Linxi Li, Yechen Wang, Liwei Jin, Carsten Maple, Ming Li
Organizations: Digital Innovation Research Center, Duke Kunshan University, Kunshan, China · OfSpectrum, Inc., Los Angeles, USA · University of Warwick, Coventry, United Kingdom · School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China
Modern audio is created by mixing stems from different sources, raising the question: can we independently watermark each stem and recover all watermarks after separation? We study a separation-first, multi-stream watermarking framework --embedding distinct information into stems using unique keys but a shared structure, mixing, separating, and decoding from each output. A naive pipeline (robust watermarking + off-the-shelf separation) yields poor bit recovery, showing robustness to generic distortions does not ensure robustness to separation artifacts. To enable this, we study separation-aware watermarking in a controlled verification pipeline, where the separator is part of the detector and can be selected or optimized together with the watermarking system. Experiments on speech+music and vocal+accompaniment mixtures show substantial gains in post-separation recovery while maintaining perceptual quality.