What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection
Authors: Jiajun Xu, Menglu Li, Xiao-Ping Zhang
Organizations: Department of Engineering Physics Tsinghua University Beijing, China · Tsinghua Shenzhen International Graduate School Tsinghua University Shenzhen, China
The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes. We therefore compare 12 acoustic representations using a shared low-capacity linear classifier to identify effective evidence under codec shift. The analysis shows that hierarchical XLS-R leads on the pooled test set, while pooled no-vocals residual statistics perform best on the unseen-codec condition, revealing complementary behavior across generation conditions. Building on this finding, we propose MN-P, a dual-view detector that integrates an utterance-level pooled no-vocals representation with token-level XLS-R features through adaptive gating. The proposed MN-P reduces EER by 54.2% overall and by 60.9% on the codec-unseen condition relative to the best-performing retrained state-of-the-art system, with consistent gains across different detector backends. These results indicate that pooled no-vocals residual statistics provide effective complementary evidence for cross-generation speech deepfake detection.
Figures & tables
Figure 1: Speech-centric single-view detection versus the proposed MN-P. (a) A conventional detector relies on a single speech representation. (b) The proposed MN-P detector complements token-level speech features with an utterance-level pooled no-vocals residual obtained after vocal separation, and integrates the two through adaptive gating.
Candidates
EER(%) ↓
CU EER(%) ↓
Source F1 ↓
B6
4.843
9.291
0.726
B2-XV
13.448
21.935
0.656
B2-NV
13.859
8.619
0.539
B2-V
25.999
41.201
0.588
B3
26.766
34.606
0.628
B1
26.956
34.204
0.620
Table 1: Six candidate representations with the lowest overall test-set EER. Source F1 is the closed-set macro-F1 of a linear probe for fake-source identity, where lower values indicate less linearly decodable source information. Bold and underlined denote the best and second-best results.
Figure 2: Architecture of the proposed MN-P detector. A frozen XLS-R branch produces a token-level speech representation, while the no-vocals branch, highlighted in the orange box, aggregates residual statistics over the whole utterance into a single vector. CrossViewGate fuses the two views before the AASIST backend. Frozen and trainable modules are marked by distinct icons.
System
AUC ↑
EER (%) ↓
Fake miss (%) ↓
CU EER (%) ↓
CU fake miss (%) ↓
MN-P-AASIST (Ours)
99.638
2.744
6.890
6.136
27.903
MN-P-Nes2Net
99.583
2.927
9.068
6.706
40.182
M0-AASIST
98.118
7.010
14.702
18.639
68.729
M0-Nes2Net
98.630
5.869
14.499
15.632
68.092
FAD [ 13 ]
98.507
6.246
14.876
16.562
70.012
W2V2-AASIST with CSAM [ 24 ]
98.647
5.986
15.444
15.692
74.325
Table 2: Performance comparison on the merged test set (%). MN-P is evaluated with AASIST and Nes2Net backends; M0 removes the pooled no-vocals view and CrossViewGate. Selected state-of-the-art systems are retrained under the same protocol. Bold and underlined denote the best and second-best results.
CU EER (%) ↓
CU fake miss (%) ↓
M0
34.830
80.180
MN-S
18.168
57.808
MN-P
8.709
14.865
Table 3: Ablation on model structure (%), on an independent train/development/test split. M0 uses XLS-R alone and MN-S feeds the no-vocals residual as a frame-level sequence. Bold and underlined denote the best and second-best results.
Audio deepfake detectors are typically evaluated on uncompressed audio, although real-world audio often undergoes low-bitrate coding. In this work, we investigate how audio coding affects deepfake detection across codecs, bitrates, and detectors, finding higher errors at lower rates. A mixed-pair protocol isolates codec-induced changes in bona fide and spoof audio, revealing asymmetric, codec-dependent failures: low-rate DAC and EnCodec mainly degrade bona fide detection, whereas X-Codec shows a stronger spoof-side limitation. Motivated by these, we propose a forensic-preserving neural audio codec (FP-NAC), which fine-tunes a pretrained codec using a detector-guided objective while preserving its native hard quantization path and bitrate. On ASVspoof 2019 LA, FP-NAC reduces EER by up to 49.8pp compared with the original DAC at 0.5kbps while maintaining comparable reconstruction quality. Although supervised by only one detector, FP-NAC improves performance across multiple detectors, highlighting forensic transparency as a codec design objective alongside perceptual quality. Our codes are available at https://github.com/kjungwoo03/FP-NAC.
Jungwoo Kim, Joonyong Park, Junyoung Koh +1
Yonsei University, Seoul, Republic of Korea · MAAP Lab, Republic of Korea · The University of Tokyo, Tokyo, Japan +1
Audio deepfakes generated by neural text-to-speech and voice-cloning systems threaten speaker verification and public discourse at scale. The core challenge is cross-dataset generalization: detectors trained on one synthesis pipeline collapse on unseen forgeries. We argue that this failure is primarily because of structural synthetic speech artifacts which are multi-timescale trajectory anomalies. Though every existing detector aggregates a fixed-window frame statistics, this misaligns the architecture with the signal. We propose FlowFake, a Liquid Time-Constant (LTC) architecture whose hidden state evolves via a learned ODE, with per-neuron adaptive time constants simultaneously resolving spectral (10ms) and prosodic (2s) cues. At only 34K parameters FlowFake achieves formal BIBO stability and O(dt^4) integration error. On a four-dataset cross domain benchmark (ASVspoof2019-LA, FakeOrReal, InTheWild, MLAAD), FlowFake reaches 75.29% on ASVspoof2019 trained only on FakeOrReal and 79.97% trained only on MLAAD. It outperforms RawGAT-ST and Whisper-DF on every evaluated pair and matching SSL Wav2vec2 (300x larger) at 0.01% of its parameter count. The source code is available on : https://github.com/GhostRider2023/FlowFake
Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, downstream finetuning has now reached a state of diminishing returns. In this paper, we shift the focus to pretraining and propose a novel recipe that combines bottleneck masked embedding prediction with flow-matching based spectrogram reconstruction. The outcome, Alethia, is the first foundational audio encoder for various voice deepfake detection and localization tasks. We evaluate on 5 different tasks with 56 benchmark datasets, and note Alethia significantly outperforms state-of-the-art SFMs with superior robustness to real-world perturbations and zero-shot generalization to unseen domains (e.g., singing deepfakes). We also demonstrate the limitation of discrete targets in masked token prediction, and show the importance of continuous embedding prediction and generative pretraining for capturing deepfake artifacts.