Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness
Authors: Zi Hu, Houmin Sun, Linxi Li, Yechen Wang, Liwei Jin, Carsten Maple, Ming Li
Organizations: Digital Innovation Research Center, Duke Kunshan University, Kunshan, China · OfSpectrum, Inc., Los Angeles, USA · University of Warwick, Coventry, United Kingdom · School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China
Neural audio codecs are challenging transformations for audio watermarking because they re-encode, quantize, and resynthesize speech. This paper investigates continuous latent-space watermarking for codec robustness. Instead of adding a watermark only to the waveform or spectrogram, we embed a 32-bit message into the continuous latent representation of a codec-like speech autoencoder. The pipeline uses a SEANet-style encoder-decoder, a Conformer-based message embedder, RVQ-guided latent decomposition, and a latent-domain detector trained under signal-processing and neural-codec transformations. Rather than proposing a final universal watermarking baseline, we characterize the trade-offs that appear when the watermark carrier is moved before neural decoding. On 48 kHz speech, EnCodec-aware training improves EnCodec-24k bit accuracy from 78.8% to 95.6% and 97.1%, while PESQ decreases from 3.727 to 3.514 and 3.427.