Modern neural audio watermarking systems typically embed a message repeatedly across time and then collapse the resulting temporal evidence into a single payload using averaging, voting, or another fixed aggregation rule. We argue that this temporal collapse limits both robustness and the recovery of multiple payloads, and that the limitation can be addressed without retraining the underlying watermarker. We freeze a pretrained watermarker's encoder and detector and train only a low-latency Conformer-based decoder. The decoder consumes the detector's temporal soft outputs, which a system-specific adapter pools into a sequence of window-level representations, and predicts the embedded message. On three frozen watermarkers (AURA, AudioSeal, and WavMark), the learned decoder improves recovery of attacked messages and yields higher detection AUROC point estimates on all three. Under controlled full- and partial-coverage multiplexing, it improves joint-exact recovery of two alternating payload words by 9.7-48.0, 6.9-17.3, and 8.2-14.2 percentage points, respectively, under one to three chained attacks on feasible clips.
Figures & tables
Attack
Operation and sampled setting
White noise
White Gaussian noise at SNR s∼Unif{1,…,15} dB.
Pink noise
1/f noise at waveform std σ=0.01 .
Low-pass
Attenuate above fc∼U(3,6) kHz.
High-pass
Attenuate below fc=500 Hz.
Smoothing
Moving average, w∼Unif{2,…,10} samples.
Resample
To f∈{44.1,24,22.05,16}∖{fop} kHz, then back.
Table 1: Signal-level augmentations. Parameters are sampled independently for each application. Unif{⋅} denotes discrete uniform sampling and U(a,b) continuous uniform sampling. † Applied only in AURA’s per-attack evaluation.
AudioSeal ( N+=2,527 )
WavMark ( N+=2,537 )
AURA ( N+=3,552 )
Attack
Base.
Conf.
Δ [95% CI]
Base.
Conf.
Δ [95% CI]
Base.
Conf.
Δ [95% CI]
White noise
6.1
6.1
0.0 [ −0.9 , +0.8 ]
0.3
0.5
+0.2 [ 0.0 , +0.4 ]
0.3
0.9
+0.5∗ [ +0.3 , +0.8 ]
Pink noise
100.0
99.9
0.0 [ −0.1 , 0.0 ]
100.0
100.0
0.0 [ 0.0 , 0.0 ]
88.3
95.2
+6.9∗ [ +5.9 , +7.9 ]
Low-pass
6.8
30.0
+23.2∗ [ +21.3 , +25.0 ]
98.6
100.0
+1.4∗ [ +0.9 , +1.9 ]
65.7
71.6
+5.8∗ [ +4.4 , +7.2 ]
High-pass
99.9
99.9
0.0 [ −0.1 , 0.0 ]
100.0
100.0
0.0 [ 0.0 , 0.0 ]
95.1
98.8
+3.7∗ [ +3.0 , +4.4 ]
Smoothing
99.8
99.8
0.0 [ −0.2 , +0.2 ]
84.7
92.9
+8.2∗ [ +7.1 , +9.3 ]
86.6
88.9
+2.3∗ [ +1.2 , +3.4 ]
Table 2: Per-attack full-coverage single-message recovery rates (%) on known-positive clips. AudioSeal and WavMark recover one 16-bit payload; AURA recovers one raw 32-bit payload without BCH. All metrics are ungated. Base. is each system’s fixed decoder (native for AudioSeal and WavMark, a BCH-free hard vote for AURA) and Conf. the Conformer; Δ=Conf.−Base. Dashes: attack not evaluated for the 16-kHz systems. ∗ denotes Holm–Bonferroni-adjusted p<0.05 within each system’s forced-attack family.
Exact message (%)
AUROC (%)
System
Regime
Conf.
Base.
Δ
Conf.
Base.
AudioSeal
Full K=2
58.6
54.9
+3.7
99.6
97.7
Full K=3
42.9
38.8
+4.2
98.8
95.5
Partial K=2
51.6
50.0
+1.6
98.2
94.8
Partial K=3
34.5
33.1
+1.4
96.5
90.9
WavMark
Full K=2
66.0
58.0
0 +8.0
97.0
96.2
Table 3: Combined-attack robustness: ungated exact-message accuracy and threshold-free detection AUROC. N/N+ is 4,992/2,527 (AudioSeal), 5,008/2,537 (WavMark), and 5,008/3,552 (AURA); AUROC uses N and exact-message accuracy uses N+ .
Full coverage
Partial coverage
System
Attacks
Baseline
Conf.
Δ [95% CI]
Baseline
Conf.
Δ [95% CI]
AURA
K=1
67.4
77.1
+9.7∗ [ +8.0,+11.4 ]
14.7
62.7
+48.0∗ [ +45.2,+50.8 ]
K=2
44.3
56.4
+12.1∗ [ +10.3,+13.9 ]
9.7
42.1
+32.4∗ [ +29.7,+35.0 ]
K=3
23.1
33.5
+10.4∗ [ +8.7,+12.2 ]
4.6
23.0
+18.4∗ [ +16.2,+20.7 ]
AudioSeal
K=1
0.1
16.8
+16.7∗ [ +14.8,+18.6 ]
1.0
18.3
+17.3∗ [ +15.1,+19.4 ]
K=2
0.1
12.8
+12.6∗ [ +11.0,+14.4 ]
0.6
13.3
+12.7∗ [ +10.8,+14.5 ]
Table 4: Controlled ABAB multiplex recovery (%). AURA is scored on two raw 32-bit words (no BCH), AudioSeal and WavMark on two 16-bit payloads; only within-system differences are interpreted. Δ is computed before rounding; ∗ : Holm-adjusted p<0.05 over 18 comparisons.
Variant
All
Partial
K=2
K=3
attacked
attacked
(full + partial)
Conformer
58.52
52.66
57.34
39.75
BiGRU
53.86
48.46
51.51
34.23
Conf.–BiGRU
+4.67
+4.20
+5.83
+5.52
− window-aux
58.14
52.45
56.69
39.24
− slot flag
58.36
52.87
57.19
39.68
Table 5: WavMark multiplex ablations: ungated joint-exact recovery (%) on feasible held-out positives, pooled over attacked conditions as defined in the text. BiGRU replaces only the Conformer trunk. One training seed.
Existing waveform-domain audio watermarks are robust to many conventional distortions but can degrade substantially under neural codec resynthesis. We investigate whether continuous neural codec latents provide a more suitable embedding space using a restricted formulation built around frozen pretrained EnCodec. To test this, a feedforward embedder maps a multi-bit payload to an additive latent perturbation decoded through the unchanged codec decoder. Compared with AudioSeal and WavMark, our latent watermark formulation degrades more gradually under repeated and low-bitrate EnCodec resynthesis, transfers to unseen DAC, and retains high detection under most waveform distortions. Substantial EnCodec robustness emerges even without codec-resynthesis supervision, indicating that this behavior is inherent to our latent formulation and is further strengthened by codec-aware training. Learned perturbations are also preserved more strongly through codec cycling than equal-norm random controls, with preservation depending more on channel-specific allocation than temporal structure. End-to-end perceptual quality remains close to that of the frozen EnCodec reconstruction, indicating that much of the observed degradation originates from the codec carrier itself. Overall, these results show that continuous neural codec latents provide a promising embedding space for watermarks that remain robust to neural codec resynthesis.
Lovro Brulec, Sahil Karawade, Leonard Kinzinger
Munich Music Labs · Technical University of Munich
Audio watermarking is increasingly important for tracing generated speech. Several audio watermarking methods have been proposed to embed the watermark in various domains, such as waveform, timbre feature, or latent representations, for making the embedded watermark robust against traditional digital signal processing (DSP) attacks. On the other hand, modern neural codecs introduce a different threat from DSP attacks: they resynthesize speech through quantized acoustic representations and can remove the embedded watermark evi- dence that is not aligned with codec-preserved structure. In this paper, we propose NeuMark, a codec-latent audio watermarking framework that embeds watermark evidence into SpeechTok- enizer acoustic tokens to address this resynthesis threat. NeuMark uses cross-attention to inject a 16-bit message across residual vector quantization (RVQ) layers, distributing the watermark over codec-aligned latent structure. Experimental results show that NeuMark substantially improves robustness under neural- codec resynthesis while supporting both watermark detection and message recovery. We also analyze the trade-off between reconstruction-referenced transparency and original-referenced robustness.
Neural audio codecs are challenging transformations for audio watermarking because they re-encode, quantize, and resynthesize speech. This paper investigates continuous latent-space watermarking for codec robustness. Instead of adding a watermark only to the waveform or spectrogram, we embed a 32-bit message into the continuous latent representation of a codec-like speech autoencoder. The pipeline uses a SEANet-style encoder-decoder, a Conformer-based message embedder, RVQ-guided latent decomposition, and a latent-domain detector trained under signal-processing and neural-codec transformations. Rather than proposing a final universal watermarking baseline, we characterize the trade-offs that appear when the watermark carrier is moved before neural decoding. On 48 kHz speech, EnCodec-aware training improves EnCodec-24k bit accuracy from 78.8% to 95.6% and 97.1%, while PESQ decreases from 3.727 to 3.514 and 3.427.
Zi Hu, Houmin Sun, Linxi Li +4
Digital Innovation Research Center, Duke Kunshan University, Kunshan, China · OfSpectrum, Inc., Los Angeles, USA · University of Warwick, Coventry, United Kingdom +1