Speech watermarking offers proactive traceability for synthetic speech, yet most existing models operate only after text-to-speech (TTS) synthesis by adding a watermark perturbation to the generated waveform. This post-hoc design leaves watermarking as an external step that can be omitted or bypassed and restricts the watermark to a shallow waveform representation. We propose NeuMark-Native, a TTS-native watermarking framework for neural codec-based synthesis. It embeds payload information into every generated codec-latent layer before waveform decoding, improving watermark persistence under downstream digital signal processing (DSP) and neural codec resynthesis. NeuMark-Native keeps the pretrained TTS model and the neural codec frozen, while optimizing only the watermark modules on generated codec tokens. Experiments on two corpora under 11 DSP attacks and 9 neural-codec attacks demonstrate robust watermark detection while preserving naturalness, intelligibility, and speech quality close to synthetic speech.
Figures & tables
Figure 1: Post-hoc watermarking applied after TTS synthesis.
Figure 2: TTS-native watermarking applied during TTS synthesis.
Figure 3: Offline preparation of paired TTS-generated codec tokens and synthetic speech.
Figure 4: NeuMark-Native training on the prepared TTS-generated codec tokens.
LibriTTS test-clean
SeedTTS-en
Method
Conditions
AUC ↑
TPR .1↑
Bit ↑
AUC ↑
TPR .1↑
Bit ↑
WavMark [ 3 ]
DSP
0.993
0.931
0.954
0.989
0.925
0.948
Codec
0.570
0.095
0.542
0.572
0.092
0.543
Overall
0.812
0.572
0.778
0.810
0.568
0.775
AudioSeal [ 15 ]
DSP
0.970
0.920
0.946
0.971
0.920
0.944
Codec
0.813
0.445
0.638
0.806
0.422
0.645
Table 1: Robustness under DSP and codec conditions. AUC: detection ROC-AUC; TPR .1 : TPR at FPR ≤0.1% ; Bit: payload accuracy. DSP, Codec, and Overall report means over their respective conditions.
LibriTTS test-clean
SeedTTS-en
Method
UTMOS ↑
SSIM ↑
WER(%) ↓
PESQ ↑
STOI ↑
UTMOS ↑
SSIM ↑
WER(%) ↓
PESQ ↑
STOI ↑
Synthetic speech
3.664
0.425
35.8
—
—
3.276
0.420
36.7
—
—
WavMark
3.632
0.421
35.7
4.159
0.996
3.224
0.414
36.8
4.081
0.988
AudioSeal
3.651
0.425
35.8
4.416
0.997
3.259
0.417
36.8
4.427
0.990
TraceableSpeech
2.198
0.174
41.3
3.048
0.925
2.109
0.182
47.1
2.463
0.895
NeuMark-Native (direct)
3.638
0.410
35.7
4.360
0.990
3.244
0.403
36.9
4.161
0.979
Table 2: Speech quality. Synthetic speech is the common unwatermarked carrier; PESQ [ 21 ] and STOI [ 18 ] use it as reference.
TTS-gen.
TTS-native
LibriTTS test-clean
SeedTTS-en
Configuration
tokens
losses
AUC ↑
TPR .1↑
Bit ↑
AUC ↑
TPR .1↑
Bit ↑
NeuMark-Native (adapted)
Yes
Yes
0.981
0.808
0.922
0.977
0.733
0.907
w/o native losses
Yes
No
0.974
0.772
0.899
0.968
0.671
0.874
w/o generated tokens
No
Yes
0.979
0.822
0.924
0.974
0.789
0.924
NeuMark-Native (direct)
No
No
0.975
0.765
0.926
0.973
0.712
0.936
Table 3: NeuMark-Native ablations. TTS-native losses are UTMOS, WavLM, and ASR objectives.
Audio watermarking is increasingly important for tracing generated speech. Several audio watermarking methods have been proposed to embed the watermark in various domains, such as waveform, timbre feature, or latent representations, for making the embedded watermark robust against traditional digital signal processing (DSP) attacks. On the other hand, modern neural codecs introduce a different threat from DSP attacks: they resynthesize speech through quantized acoustic representations and can remove the embedded watermark evi- dence that is not aligned with codec-preserved structure. In this paper, we propose NeuMark, a codec-latent audio watermarking framework that embeds watermark evidence into SpeechTok- enizer acoustic tokens to address this resynthesis threat. NeuMark uses cross-attention to inject a 16-bit message across residual vector quantization (RVQ) layers, distributing the watermark over codec-aligned latent structure. Experimental results show that NeuMark substantially improves robustness under neural- codec resynthesis while supporting both watermark detection and message recovery. We also analyze the trade-off between reconstruction-referenced transparency and original-referenced robustness.
As speech generation models become increasingly realistic and widely accessible, concerns about the misuse, attribution, and governance of synthetic speech continue to grow. Watermarking provides a practical way to make synthesized speech traceable and verifiable. Most existing speech watermarking methods embed watermark information into signal-level representations, such as waveforms or spectrograms. Under sufficiently strong distortions, the embedded watermark may be weakened or destroyed, leading to degraded detectability. In this paper, we propose SSTMark, a training-free speech watermarking framework that operates at the semantic level through text watermarking. Unlike conventional signal-level watermarking methods, SSTMark encodes watermark information into the semantic content conveyed by generated speech, and detects the watermark from the recovered linguistic content. Experiments on AudioMarkBench demonstrate that SSTMark exhibits the strongest average robustness. Compared with the state-of-the-art baselines at a fixed false positive rate of 1%, SSTMark improves the average detection rate by 4.6% and 16.9% on signal-processing edits and compression edits, respectively.
Kuan-Lin Chu, Jun-Cheng Chen, Chun-Shien Lu
CITI, Academia Sinica · Taiwan, ROC · IIS, Academia Sinica
Large language model (LLM)-based text-to-speech (TTS) models have achieved remarkable voice cloning capabilities, raising concerns about potential deepfake misuse. Speech watermarking mitigates this by embedding traceable information into generated speech. Mainstream watermarking methods operate at the signal level (waveform or spectrogram), rendering the watermark vulnerable to generative attacks (e.g., neural codec and vocoder). To address this, we propose DuraMark, a robust information-level watermarking framework. It utilizes syllable duration editing to achieve watermark embedding. Specifically, DuraMark integrates a duration-controllable LLM-based TTS model to edit syllable durations during synthesis, coupled with a duration extractor to extract these durations for detection. Experiments demonstrate DuraMark's superior robustness against generative attacks, significantly outperforming signal-level baselines. Audio samples are available at https://muzw.github.io/duramark_demo/.
Zhenwei Mou, Weili Jiang, Liping Chen +4
University of Science and Technology of China, China · Institute of Forensic Science, Ministry of Public Security, China · The Hong Kong Polytechnic University, China