Watermarking is a promising tool for establishing the provenance of AI-generated speech. While many neural audio watermarking methods rely on a separately trained watermark generator, token-level watermarking is a training-free alternative that operates directly during generation. Its main weakness is retokenization: decoding generated speech to a waveform and encoding it again can change token identities and erode the watermark. To make the watermark robust to these changes, we propose Redwing, REtokenization-Durable Watermarking IN Generation. It builds a graph from the token substitutions observed under retokenization, whose Laplacian yields a basis that assigns similar values to tokens likely to substitute for one another. Over this basis, embedding and detection functions are jointly optimized to preserve watermark signal through retokenization while limiting embedding distortion and detector variability on unwatermarked speech. On the Moshi full-duplex system, after eight consecutive passes of Mimi resynthesis, Redwing achieves 80.7% TPR at a calibrated 1% FPR, compared with 8.3% for KGW and at most 7.3% for WMAR. It also has the highest TPR after eight passes through three other neural codecs (77.5-93.0%), and the gains generalize to TTS models at a speech-quality cost close to that of KGW. These results show that retokenization is not merely a source of noise: its transition structure can be exploited as a design principle for robust token-level watermarking.
Figures & tables
Method
Mimi (native)
EnCodec
SpeechTok.
SNAC
DAC16
Redwing
80.7 ± 3.2
85.0 ± 2.9
77.5 ± 3.3
93.0 ± 2.1
34.2 ± 3.8
KGW
8.3 ± 2.2
16.5 ± 3.0
5.3 ± 1.8
63.2 ± 3.8
12.5 ± 2.6
WMAR FT
5.5 ± 1.8
24.7 ± 3.4
2.2 ± 1.2
58.5 ± 3.9
10.3 ± 2.4
WMAR FT+Augs
7.3 ± 2.1
15.7 ± 2.9
2.0 ± 1.2
57.2 ± 3.9
20.7 ± 3.2
AudioSeal
22.5 ∗ ± 3.3
82.3 ∗ ± 3.0
1.5 ± 1.0
0.0 ± 0.3
0.2 ± 0.5
Timbre
0.0 ± 0.3
0.3 ± 0.6
0.0 ± 0.3
0.0 ± 0.3
0.0 ± 0.3
Table 1: Detection on Moshi after eight codec passes. Values are TPR (%) at the fixed threshold ± the half-width of the 95 % interval. Post-hoc methods are below the line, and bold is as defined in Section 4 . ∗ The same attack also flags more than 5 % of unwatermarked answers (Table 14 ), so the cell is not bold. WavMark, which detects 0 % after every codec here, is omitted (Appendix C.1 ).
Method
Native
Mimi
EnCodec
SpeechTok.
SNAC
DAC16
CosyVoice3
Redwing
84.7 ± 2.9
68.8 ± 3.7
89.6 ± 2.5
61.2 ± 3.9
91.9 ± 2.2
94.0 ± 1.9
KGW
9.2 ± 2.3
1.2 ± 0.9
9.9 ± 2.4
2.8 ± 1.4
29.3 ± 3.6
21.9 ± 3.3
AudioSeal
15.1 ∗ ± 2.9
30.0 ∗ ± 3.7
56.4 ∗ ± 4.0
0.3 ± 0.6
0.0 ± 0.3
0.2 ± 0.5
Timbre
23.1 ± 3.4
0.3 ± 0.6
3.0 ± 1.4
0.0 ± 0.3
0.0 ± 0.3
0.2 ± 0.5
CRAW
5.7 ± 1.9
1.2 ± 0.9
61.5 ± 3.9
10.7 ± 2.5
11.7 ± 2.6
69.7 ± 3.7
Table 2: Detection on the TTS models after eight codec passes. The layout is as in Table 1 , with each model’s own resynthesis (Native) as the first codec.
Method
UTMOSv2
DNSMOS Pro
NISQAv2
NISQA-TTS
KL/frame
Moshi (non-empty answers)
Unwatermarked
3.15 ± 0.03
4.38 ± 0.04
4.64 ± 0.04
3.56 ± 0.05
0
Unwatermarked, FT decoder †
3.04 ± 0.03
4.33 ± 0.04
4.55 ± 0.04
3.40 ± 0.04
0
Unwatermarked, FT+Augs decoder †
3.16 ± 0.03
4.36 ± 0.04
4.61 ± 0.04
3.54 ± 0.05
0
Redwing
3.05 ± 0.03
4.47 ± 0.02
4.62 ± 0.03
3.52 ± 0.04
1.52
KGW
3.06 ± 0.04
4.35 ± 0.05
4.51 ± 0.06
3.44 ± 0.05
1.73
Table 3: Speech quality at the selected strengths. Mean ± 95 % CI (higher is better), for Moshi on non-empty answers. KL/frame is the cost. † Decoded with WMAR’s fine-tuned decoder, which changes quality even without a watermark. Bold (Section 4 ) covers only the watermarked rows.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Stream
1
2
3
4
5
6
7
8
Survival share
0.72
0.39
0.26
0.25
0.20
0.21
0.19
0.17
Participation, without
116
161
228
172
99
183
70
183
Participation, with
79
151
1.2
1.1
1.1
1.2
1.2
1.0
Single-token modes, with
0
2
15
14
15
11
14
15
Appendix
Table 4: Self-transitions and the basis on Moshi. The counts come from Moshi’s own Mimi channel, on the support of tokens with at least 50 occurrences. The survival share is the fraction of counted transitions that are self-transitions. Participation is the median participation ratio, the effective number of tokens covered, of the 16 leading nontrivial modes of the normalized adjacency, computed without and with self-transitions. The last row counts how many of these 16 modes put at least half of their mass on one token when self-transitions are kept.
Functions
δ
KL (nats/frame)
UTMOSv2 difference
TPR (%)
Unbounded
0.05
0.05
−0.04
2
Unbounded
0.06
0.24
−0.05
7
Unbounded
0.07
0.94
−0.48
35
Unbounded
0.08
1.94
−1.00
68
Bounded ( κ=5 )
0.70
1.54
+0.03
97
Appendix
Table 5: Unbounded and bounded functions. All values are for Moshi’s streams 1–4 on the validation set (250 prompts). KL is the divergence per frame, summed over the four streams. The UTMOSv2 difference is paired against unwatermarked speech, and the TPR is at the fixed threshold without attack.
Stream
κ
1
2
3
4
5
6
7
8
Mean
2
0.31
0.26
0.41
0.70
0.57
0.58
0.64
0.63
0.51
3
0.49
0.49
0.79
0.94
0.78
0.86
0.84
0.85
0.76
4
0.65
0.65
0.89
0.98
0.79
0.96
0.85
0.90
0.83
5
0.82
0.68
0.90
1.00
0.79
1.00
0.87
0.93
0.87
7
0.91
0.69
0.90
1.00
0.80
1.00
0.90
0.96
0.90
Appendix
Table 6: Objective retained under the bound κ . Each value is the attained objective divided by the unconstrained optimum σ1(M) , for one Moshi stream.
Model
Clips
τ⋆=−2
−1
0
+1
+2
Moshi
watermarked (600)
0.000
0.985
0.000
0.000
0.000
Moshi
unwatermarked (600)
0.055
0.073
0.047
0.058
0.070
CosyVoice3
watermarked (596)
0.002
0.000
0.988
0.002
0.002
CosyVoice3
unwatermarked (597)
0.055
0.047
0.062
0.049
0.060
MOSS-TTS
watermarked (593)
0.000
0.000
1.000
0.000
0.000
MOSS-TTS
unwatermarked (599)
0.048
0.047
0.078
0.053
0.060
Appendix
Table 7: Where the evidence peaks. Each value is the fraction of test clips whose maximizing offset τ⋆ takes the given value, when the search covers ±8 frames and no attack is applied. An even spread would put 1/17, or 0.059, on each offset.
Set
Used for
Moshi
CosyVoice3
MOSS-TTS
Substitution counts
graph, basis and P
20 hours of LibriSpeech
Validation
strength δ and watermarked streams
250
250
250
Calibration
thresholds
1,200
1,000 (994)
1,000 (993)
Test
all reported results
600
600 (596–598)
600 (593–599)
Appendix
Table 8: Data sets. The numbers count prompts for Moshi and texts for the TTS models, and the calibration sets count unwatermarked answers. On the TTS models, a few generations failed or exceeded the tokenizer’s length limit and are excluded, and parentheses give the clips left.
Codec
Implementation
Setting
Mimi
Mimi of Moshi ( kyutai/moshiko-pytorch-bf16 )
8 quantizers
CosyVoice3 native
speech tokenizer, flow-matching decoder and HiFT vocoder
clip as its own prompt
MOSS-TTS native
MOSS audio tokenizer
32 quantizers
EnCodec
facebook/encodec_24khz
6 kbps
SpeechTokenizer
speechtokenizer_hubert_avg
all quantizers, 24 kHz input
SNAC
hubertsiuzdak/snac_24khz
all quantizers
Appendix
Table 9: Codecs. The table lists the checkpoint and the setting of each codec used in the attacks.
Set
Method
FPR
Identity
Mimi ×1
Mimi ×8
EnCodec ×8
Test
Aligned-IS, h=20
1.7
4.7
3.3
2.0
0.5
Test
KGW
1.5
78.5
64.5
8.3
16.5
Validation
Aligned-IS, h=20
2.0
4.4
4.0
2.4
0.4
Validation
Aligned-IS, h=50
1.2
32.0
9.2
2.4
1.6
Validation
Aligned-IS, h=100
0.8
50.0
14.0
1.2
6.4
Validation
streams 1–4 only
1.2
3.2
2.0
2.0
0.4
Appendix
Table 10: Aligned-IS on Moshi. TPR (%) at the fixed threshold without attack (Identity), after one and eight Mimi passes, and after eight EnCodec passes. FPR is the rate on the unwatermarked answers of the same set without attack. The indented variants use h=20 , and KGW is shown for reference.
Method
DAC24
DAC44
Opus
MP3
AAC
Moshi
Redwing
64.0 ± 3.8
94.8 ± 1.8
90.7 ± 2.3
93.5 ± 2.0
94.3 ± 1.9
KGW
36.3 ± 3.8
74.5 ± 3.5
57.8 ± 3.9
53.2 ± 4.0
73.8 ± 3.5
WMAR FT
30.0 ± 3.7
75.3 ± 3.4
61.0 ± 3.9
57.8 ± 3.9
76.7 ± 3.4
WMAR FT+Augs
57.2 ± 3.9
75.0 ± 3.5
52.7 ± 4.0
61.8 ± 3.9
76.3 ± 3.4
AudioSeal
100.0 ± 0.3
100.0 ± 0.3
76.3 ± 3.4
100.0 ± 0.3
100.0 ± 0.3
Appendix
Table 11: Detection after eight passes through high-fidelity codecs. TPR (%) on the test sets at the fixed threshold, ± the half-width of the Wilson 95 % interval. DAC24 and DAC44 are the 24 and 44.1 kHz DAC models. Opus runs at 24 kbps, and MP3 and AAC at 64 kbps. Bold is as defined in Section 4 , within each model. None of these attacks flags more than 5 % of unwatermarked answers, so no cell is excluded from the bold as in Table 1 .
Moshi
CosyVoice3
MOSS-TTS
Attack
Redwing
KGW
WMAR FT
WMAR FT+Augs
AudioSeal
WavMark
Timbre
CRAW
Redwing
KGW
AudioSeal
WavMark
Timbre
CRAW
Redwing
KGW
AudioSeal
WavMark
Timbre
CRAW
None
95
78
84
82
100
89
100
100
98
92
100
100
100
100
100
98
100
100
100
100
White noise (SNR)
40 dB
95
79
76
82
100
88
100
100
98
91
100
97
99
100
100
83
100
95
99
100
30 dB
95
78
45
81
100
75
89
100
98
89
100
61
86
100
100
53
100
66
85
100
20 dB
87
47
1
80
100
3
31
99
97
84
98
4
39
100
100
14
97
7
40
99
Appendix
Table 12: Detection after signal-processing attacks. TPR (%) on the test sets at the fixed threshold, rounded to whole percent. The Wilson 95 % interval of every value is at most 8 percentage points wide. Within each row and model, bold marks the highest TPR and every TPR whose interval overlaps it, and nothing when all overlap. ∗ The same attack also flags more than 5 % of unwatermarked answers, so the cell is excluded from the bold. WavMark does not embed into 67 of Moshi’s answers, which caps its TPR there at 89 %.
After eight passes
Method
Identity
Mimi
EnCodec
SpeechTok.
SNAC
DAC16
DAC24
DAC44
Opus
MP3
AAC
Redwing
0.5
0.7
0.5
0.1
0.6
0.3
0.8
0.6
0.6
0.7
0.7
KGW
0.7
0.8
0.5
0.2
0.7
0.4
1.0
0.5
1.0
0.7
0.4
AudioSeal
4.3
42.2 ∗
79.1 ∗
2.0
2.7
4.5
3.5
4.3
7.0 ∗
4.4
3.4
WavMark
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
Timbre
0.7
0.0
1.6
0.1
0.0
0.1
0.0
0.0
0.1
0.5
1.2
Appendix
Table 13: False positives on human speech. FPR (%) on the 1,000 WildVoice recordings that serve as Moshi’s prompts, at each method’s fixed threshold for Moshi, without attack (Identity) and after eight passes of each codec. ∗ Values above 5 %.
After eight passes
Method
Identity
Native
Mimi
EnCodec
SpeechTok.
SNAC
DAC16
DAC24
DAC44
Opus
MP3
AAC
Moshi
Redwing
1.0
0.5
–
0.3
0.3
0.5
0.0
0.3
0.8
0.7
0.2
0.8
KGW
1.5
0.3
–
1.7
0.3
1.2
0.5
1.2
2.5
0.5
1.2
1.0
WMAR FT
1.3
0.0
–
0.5
0.3
0.3
0.2
0.5
0.5
0.3
0.2
0.3
WMAR FT+Augs
1.7
0.5
–
1.3
0.7
1.7
1.5
1.8
1.2
1.2
1.0
1.8
Appendix
Table 14: False positives after attack. FPR (%) on the unwatermarked test answers at each method’s fixed threshold, without attack (Identity) and after eight passes. Native is each model’s own codec, which is Mimi on Moshi. ∗ Values above 5 %.
Method
38–50
51–100
101–150
151–200
All
No attack
Redwing
75.6 ± 7.6
98.8 ± 2.0
100.0 ± 1.3
100.0 ± 1.1
94.8 ± 1.8
KGW
30.0 ± 6.8
93.6 ± 4.1
99.0 ± 2.5
100.0 ± 1.0
78.5 ± 3.3
After eight Mimi passes
Redwing
39.5 ± 8.7
75.8 ± 6.6
96.5 ± 3.3
100.0 ± 1.1
80.7 ± 3.2
KGW
1.8 ± 2.2
8.5 ± 4.7
9.6 ± 5.7
13.5 ± 4.9
8.3 ± 2.2
Appendix
Table 15: Detection on Moshi by answer length. TPR (%) at the fixed threshold on the test set ± the 95 % Wilson half-width, by answer length in frames. Each bin covers 50 frames, or 4 s, and All is the TPR on all test answers.
Method
UTMOSv2
DNSMOS Pro
NISQAv2
NISQA-TTS
Moshi (non-empty answers)
Unwatermarked
3.15 ± 0.03
4.38 ± 0.04
4.64 ± 0.04
3.56 ± 0.05
AudioSeal
3.14 ± 0.03
4.37 ± 0.04
4.50 ± 0.04
3.32 ± 0.05
WavMark
2.88 ± 0.04
4.26 ± 0.04
4.54 ± 0.04
3.39 ± 0.04
Timbre
3.11 ± 0.03
4.31 ± 0.04
4.55 ± 0.05
3.30 ± 0.04
CRAW
2.92 ± 0.04
4.25 ± 0.05
3.96 ± 0.07
3.57 ± 0.05
Appendix
Table 16: Speech quality of the post-hoc methods. Mean ± 95 % CI on the test sets, as in Table 3 . Bold as in Table 3 .
Method
Δ UTMOSv2
Δ DNSMOS Pro
Δ NISQAv2
Δ NISQA-TTS
Moshi (prompts non-empty for all)
Redwing
−0.09± 0.04
+0.10± 0.05
0.00± 0.06
−0.03± 0.06
KGW
−0.09± 0.06
−0.04± 0.07
−0.15± 0.08
−0.14± 0.07
WMAR FT †
−0.07± 0.05
−0.04± 0.07
−0.12± 0.07
−0.16± 0.07
WMAR FT+Augs †
−0.11± 0.05
−0.05± 0.07
−0.15± 0.08
−0.15± 0.07
AudioSeal
−0.01± 0.02
−0.01± 0.01
−0.15± 0.01
−0.24± 0.02
Appendix
Table 17: Paired differences in speech quality. Mean difference from the unwatermarked answer to the same prompt ± the 95 % CI on the test sets. On Moshi, only prompts with non-empty answers from the unwatermarked model, ours and KGW are used, and on the TTS models all clips. † Compared with the unwatermarked answers decoded by the same fine-tuned decoder.
After eight passes
Variant
δ
KL
UTMOSv2
Identity
Mimi
EnCodec
SpeechTok.
DAC16
Unwatermarked
–
0
3.15 ± 0.03
–
–
–
–
–
KGW
2
1.73
3.06 ± 0.04
78.5 ± 3.3
8.3 ± 2.2
16.5 ± 3.0
5.3 ± 1.8
12.5 ± 2.6
Redwing
0.7
1.52
3.05 ± 0.03
94.8 ± 1.8
80.7 ± 3.2
85.0 ± 2.9
77.5 ± 3.3
34.2 ± 3.8
Basis
Random basis
0.7
1.60
3.12 ± 0.03
92.0 ± 2.2
52.3 ± 4.0
36.0 ± 3.8
24.5 ± 3.4
45.5 ± 4.0
Appendix
Table 18: Ablations on Moshi. TPR (%) at the fixed threshold on the test set ± the 95 % Wilson half-width, and UTMOSv2 on non-empty answers ± the 95 % CI. Each variant removes one component and runs at its own budget-matched δ , and KL is its cost per frame. The random basis uses 16 random orthonormal directions, and the transition basis the leading singular vectors of the transition matrix P . The random function is one fixed random g=h per stream, without basis or solve. ‡ The same watermarked answers as ours, but the detector scores them with the embedding function g in place of the solved detection function h , which tests whether a separate detection function is needed. Bold is as defined in Section 4 .
Recent advances in generative speech models have made it increasingly difficult to distinguish authentic from synthetic audio, enabling new forms of fraud and misinformation. Audio watermarking offers a promising defense by embedding an imperceptible signal into generated speech that can later be detected to verify its provenance. However, recent studies have shown that existing post-hoc watermarking methods fail under neural codecs and denoisers, transformations routinely applied during real-world storage, transmission, and processing, severely limiting their practical utility. Here we introduce CRAW, a codec-robust audio watermarking framework that jointly improves robustness against neural re-synthesis while maintaining high perceptual quality. CRAW combines distortion-aware training with an attention-based pooling mechanism, inference-time perceptual mask- ing, and an error-correcting code to recover the fidelity lost during robust training. Experiments demonstrate that CRAW achieves state-of-the-art robustness against neural codecs, denoisers, and vocoders while maintaining perceptual quality comparable to existing post-hoc watermarking methods. The code is available at https://github.com/DavidC1212/craw.
As policy catches up with the capabilities of generative AI, watermarking is central to content provenance efforts. Inference-time watermarks for autoregressive models are unfit for continuous modalities due to discretization inconsistencies. Existing methods overcome this by finetuning the modality tokenizers, nullifying the watermark's training-free advantage. In this work, motivated by the vocabulary redundancy of discretization, we propose an elegant solution for powerful and robust watermarking of synthetic audio. We theoretically analyze the impact of token errors on watermark detection, and effectively mitigate them using a reduced vocabulary obtained via community detection. Thorough experiments showcase that our gradient-free method can boost detectability by several orders of magnitude, while also achieving built-in robustness to audio modifications. Broadly, we discover a new state-of-the-art for token-level watermarks in multimedia, which simply arises from the nature of discrete representation learning.
Georgios Milis, Yubin Qin, Yihan Wu +1
Department of Computer Science, University of Maryland, College Park, USA.
Neural audio watermarks are increasingly deployed in commercial speech generation systems to make AI-generated speech traceable, yet their robustness has been studied mainly under conventional signal distortions. Since a watermark can be regarded as imperceptible noise added to the speech signal, a natural question is whether speech enhancement (SE), as a denoising model, can remove it. In this paper, we cascade Gaussian noise with SE models as a black-box watermark removal attack, covering both discriminative and generative SE paradigms, against six neural watermarks: AudioSeal, WavMark, SilentCipher, Timbre, Perth, and AlignMark. Experimental results show that the proposed attack significantly outperforms existing neural re-synthesis methods in watermark removal. In particular, we find that generative SE, which reconstructs the harmonic regions of speech while denoising, is highly destructive to watermarks. These findings show that SE poses a serious threat to current audio watermarking methods, and we call for SE-aware robustness evaluation in watermark design.
Xincong Zhong, Shengyao Wang, Lingfeng Yao +4
Waseda University, Japan · University of Houston, USA