Modern neural audio watermarking systems typically embed a message repeatedly across time and then collapse the resulting temporal evidence into a single payload using averaging, voting, or another fixed aggregation rule. We argue that this temporal collapse limits both robustness and the recovery of multiple payloads, and that the limitation can be addressed without retraining the underlying watermarker. We freeze a pretrained watermarker's encoder and detector and train only a low-latency Conformer-based decoder. The decoder consumes the detector's temporal soft outputs, which a system-specific adapter pools into a sequence of window-level representations, and predicts the embedded message. On three frozen watermarkers (AURA, AudioSeal, and WavMark), the learned decoder improves recovery of attacked messages and yields higher detection AUROC point estimates on all three. Under controlled full- and partial-coverage multiplexing, it improves joint-exact recovery of two alternating payload words by 9.7-48.0, 6.9-17.3, and 8.2-14.2 percentage points, respectively, under one to three chained attacks on feasible clips.
Figures & tables
Attack
Operation and sampled setting
White noise
White Gaussian noise at SNR s∼Unif{1,…,15} dB.
Pink noise
1/f noise at waveform std σ=0.01 .
Low-pass
Attenuate above fc∼U(3,6) kHz.
High-pass
Attenuate below fc=500 Hz.
Smoothing
Moving average, w∼Unif{2,…,10} samples.
Resample
To f∈{44.1,24,22.05,16}∖{fop} kHz, then back.
Table 1: Signal-level augmentations. Parameters are sampled independently for each application. Unif{⋅} denotes discrete uniform sampling and U(a,b) continuous uniform sampling. † Applied only in AURA’s per-attack evaluation.
AudioSeal ( N+=2,527 )
WavMark ( N+=2,537 )
AURA ( N+=3,552 )
Attack
Base.
Conf.
Δ [95% CI]
Base.
Conf.
Δ [95% CI]
Base.
Conf.
Δ [95% CI]
White noise
6.1
6.1
0.0 [ −0.9 , +0.8 ]
0.3
0.5
+0.2 [ 0.0 , +0.4 ]
0.3
0.9
+0.5∗ [ +0.3 , +0.8 ]
Pink noise
100.0
99.9
0.0 [ −0.1 , 0.0 ]
100.0
100.0
0.0 [ 0.0 , 0.0 ]
88.3
95.2
+6.9∗ [ +5.9 , +7.9 ]
Low-pass
6.8
30.0
+23.2∗ [ +21.3 , +25.0 ]
98.6
100.0
+1.4∗ [ +0.9 , +1.9 ]
65.7
71.6
+5.8∗ [ +4.4 , +7.2 ]
High-pass
99.9
99.9
0.0 [ −0.1 , 0.0 ]
100.0
100.0
0.0 [ 0.0 , 0.0 ]
95.1
98.8
+3.7∗ [ +3.0 , +4.4 ]
Smoothing
99.8
99.8
0.0 [ −0.2 , +0.2 ]
84.7
92.9
+8.2∗ [ +7.1 , +9.3 ]
86.6
88.9
+2.3∗ [ +1.2 , +3.4 ]
Table 2: Per-attack full-coverage single-message recovery rates (%) on known-positive clips. AudioSeal and WavMark recover one 16-bit payload; AURA recovers one raw 32-bit payload without BCH. All metrics are ungated. Base. is each system’s fixed decoder (native for AudioSeal and WavMark, a BCH-free hard vote for AURA) and Conf. the Conformer; Δ=Conf.−Base. Dashes: attack not evaluated for the 16-kHz systems. ∗ denotes Holm–Bonferroni-adjusted p<0.05 within each system’s forced-attack family.
Exact message (%)
AUROC (%)
System
Regime
Conf.
Base.
Δ
Conf.
Base.
AudioSeal
Full K=2
58.6
54.9
+3.7
99.6
97.7
Full K=3
42.9
38.8
+4.2
98.8
95.5
Partial K=2
51.6
50.0
+1.6
98.2
94.8
Partial K=3
34.5
33.1
+1.4
96.5
90.9
WavMark
Full K=2
66.0
58.0
0 +8.0
97.0
96.2
Table 3: Combined-attack robustness: ungated exact-message accuracy and threshold-free detection AUROC. N/N+ is 4,992/2,527 (AudioSeal), 5,008/2,537 (WavMark), and 5,008/3,552 (AURA); AUROC uses N and exact-message accuracy uses N+ .
Full coverage
Partial coverage
System
Attacks
Baseline
Conf.
Δ [95% CI]
Baseline
Conf.
Δ [95% CI]
AURA
K=1
67.4
77.1
+9.7∗ [ +8.0,+11.4 ]
14.7
62.7
+48.0∗ [ +45.2,+50.8 ]
K=2
44.3
56.4
+12.1∗ [ +10.3,+13.9 ]
9.7
42.1
+32.4∗ [ +29.7,+35.0 ]
K=3
23.1
33.5
+10.4∗ [ +8.7,+12.2 ]
4.6
23.0
+18.4∗ [ +16.2,+20.7 ]
AudioSeal
K=1
0.1
16.8
+16.7∗ [ +14.8,+18.6 ]
1.0
18.3
+17.3∗ [ +15.1,+19.4 ]
K=2
0.1
12.8
+12.6∗ [ +11.0,+14.4 ]
0.6
13.3
+12.7∗ [ +10.8,+14.5 ]
Table 4: Controlled ABAB multiplex recovery (%). AURA is scored on two raw 32-bit words (no BCH), AudioSeal and WavMark on two 16-bit payloads; only within-system differences are interpreted. Δ is computed before rounding; ∗ : Holm-adjusted p<0.05 over 18 comparisons.
Variant
All
Partial
K=2
K=3
attacked
attacked
(full + partial)
Conformer
58.52
52.66
57.34
39.75
BiGRU
53.86
48.46
51.51
34.23
Conf.–BiGRU
+4.67
+4.20
+5.83
+5.52
− window-aux
58.14
52.45
56.69
39.24
− slot flag
58.36
52.87
57.19
39.68
Table 5: WavMark multiplex ablations: ungated joint-exact recovery (%) on feasible held-out positives, pooled over attacked conditions as defined in the text. BiGRU replaces only the Conformer trunk. One training seed.
Digital Innovation Research Center, Duke Kunshan University, Kunshan, China · OfSpectrum, Inc., Los Angeles, USA · University of Warwick, Coventry, United Kingdom +1