cs.SDOct 4, 2026

NeuMark-Native: Robust Text-to-Speech-Native Watermarking Through Full Utilization of Neural Audio Codec Latent Space

Authors: Annan Wu, Wen-Chin Huang, Tomoki Toda

Organizations: Nagoya University, Japan

Abstract

Speech watermarking offers proactive traceability for synthetic speech, yet most existing models operate only after text-to-speech (TTS) synthesis by adding a watermark perturbation to the generated waveform. This post-hoc design leaves watermarking as an external step that can be omitted or bypassed and restricts the watermark to a shallow waveform representation. We propose NeuMark-Native, a TTS-native watermarking framework for neural codec-based synthesis. It embeds payload information into every generated codec-latent layer before waveform decoding, improving watermark persistence under downstream digital signal processing (DSP) and neural codec resynthesis. NeuMark-Native keeps the pretrained TTS model and the neural codec frozen, while optimizing only the watermark modules on generated codec tokens. Experiments on two corpora under 11 DSP attacks and 9 neural-codec attacks demonstrate robust watermark detection while preserving naturalness, intelligibility, and speech quality close to synthetic speech.

Figures & tables

Explore similar work

CardsList
  1. SSTMark: Robust Training-Free Semantic-Level Speech Watermarking

    Jul 20, 2026Kuan-Lin Chu, Jun-Cheng Chen, Chun-Shien LuAudio WatermarkingRobust Watermarking

  2. DuraMark: Duration-Embedded Watermarking in LLM-based TTS

    Jun 13, 2026Zhenwei Mou, Weili Jiang, Liping Chen +4Audio WatermarkingTTS Synthesis