Speech watermarking offers proactive traceability for synthetic speech, yet most existing models operate only after text-to-speech (TTS) synthesis by adding a watermark perturbation to the generated waveform. This post-hoc design leaves watermarking as an external step that can be omitted or bypassed and restricts the watermark to a shallow waveform representation. We propose NeuMark-Native, a TTS-native watermarking framework for neural codec-based synthesis. It embeds payload information into every generated codec-latent layer before waveform decoding, improving watermark persistence under downstream digital signal processing (DSP) and neural codec resynthesis. NeuMark-Native keeps the pretrained TTS model and the neural codec frozen, while optimizing only the watermark modules on generated codec tokens. Experiments on two corpora under 11 DSP attacks and 9 neural-codec attacks demonstrate robust watermark detection while preserving naturalness, intelligibility, and speech quality close to synthetic speech.
Figures & tables
Figure 1: Post-hoc watermarking applied after TTS synthesis.
Figure 2: TTS-native watermarking applied during TTS synthesis.
Figure 3: Offline preparation of paired TTS-generated codec tokens and synthetic speech.
Figure 4: NeuMark-Native training on the prepared TTS-generated codec tokens.
LibriTTS test-clean
SeedTTS-en
Method
Conditions
AUC ↑
TPR .1↑
Bit ↑
AUC ↑
TPR .1↑
Bit ↑
WavMark [ 3 ]
DSP
0.993
0.931
0.954
0.989
0.925
0.948
Codec
0.570
0.095
0.542
0.572
0.092
0.543
Overall
0.812
0.572
0.778
0.810
0.568
0.775
AudioSeal [ 15 ]
DSP
0.970
0.920
0.946
0.971
0.920
0.944
Codec
0.813
0.445
0.638
0.806
0.422
0.645
Table 1: Robustness under DSP and codec conditions. AUC: detection ROC-AUC; TPR .1 : TPR at FPR ≤0.1% ; Bit: payload accuracy. DSP, Codec, and Overall report means over their respective conditions.
LibriTTS test-clean
SeedTTS-en
Method
UTMOS ↑
SSIM ↑
WER(%) ↓
PESQ ↑
STOI ↑
UTMOS ↑
SSIM ↑
WER(%) ↓
PESQ ↑
STOI ↑
Synthetic speech
3.664
0.425
35.8
—
—
3.276
0.420
36.7
—
—
WavMark
3.632
0.421
35.7
4.159
0.996
3.224
0.414
36.8
4.081
0.988
AudioSeal
3.651
0.425
35.8
4.416
0.997
3.259
0.417
36.8
4.427
0.990
TraceableSpeech
2.198
0.174
41.3
3.048
0.925
2.109
0.182
47.1
2.463
0.895
NeuMark-Native (direct)
3.638
0.410
35.7
4.360
0.990
3.244
0.403
36.9
4.161
0.979
Table 2: Speech quality. Synthetic speech is the common unwatermarked carrier; PESQ [ 21 ] and STOI [ 18 ] use it as reference.
TTS-gen.
TTS-native
LibriTTS test-clean
SeedTTS-en
Configuration
tokens
losses
AUC ↑
TPR .1↑
Bit ↑
AUC ↑
TPR .1↑
Bit ↑
NeuMark-Native (adapted)
Yes
Yes
0.981
0.808
0.922
0.977
0.733
0.907
w/o native losses
Yes
No
0.974
0.772
0.899
0.968
0.671
0.874
w/o generated tokens
No
Yes
0.979
0.822
0.924
0.974
0.789
0.924
NeuMark-Native (direct)
No
No
0.975
0.765
0.926
0.973
0.712
0.936
Table 3: NeuMark-Native ablations. TTS-native losses are UTMOS, WavLM, and ASR objectives.
University of Science and Technology of China, China · Institute of Forensic Science, Ministry of Public Security, China · The Hong Kong Polytechnic University, China