Exposing and Mitigating Neural Codec Vulnerabilities in Audio Deepfake Detection
Organizations: College of Innovation and Technology University of Michigan-Flint Michigan, USA
Abstract
Existing audio deepfake detection (ADD) datasets and detectors are primarily built for vocoder-based synthesis, evaluated against traditional post-hoc perturbations such as MP3/AAC compression or additive noise, applied independently of generation. However, recent speech synthesizers, particularly ALM-based systems, use neural audio codecs both for compression and as the resynthesis reconstructing waveforms from generated tokens, producing artifacts distinct from post-hoc compression. Neural codecs thus play a dual role: some are designed for pure compression under low-bandwidth communication, while others serve as resynthesis components. Despite this dual role, robustness to codec-based compression, unlike post-hoc compression, remains largely unexplored. We expose this gap, showing that state-of-the-art (SOTA) ADD models degrade drastically on codec-compressed speech; in particular, systems trained on Codec Resynthesized data as a proxy for codec-based generation prove most vulnerable, with legitimately compressed bona fide speech often misclassified as fake. To investigate this, we construct the Audio Neural Codec-Spoof dataset by applying seven neural codec algorithms to existing ADD benchmarks, isolating codec-induced resynthesis artifacts as a controlled proxy for codec-based generation. As baseline mitigation, we propose PCL-NET (Pairwise Consistency Learned Network), fine-tuning a pretrained XLS-R (300M) encoder with a pairwise consistency objective that minimizes the representation distance between an utterance's uncompressed and codec-compressed versions, disentangling codec artifacts from the real-versus-fake decision. As a result, PCL-NET reduces average EER under neural codec compression from 28.67% to 12.77%, while preserving competitive CoSG-based deepfake detection performance. We will also make the dataset publicly available on Hugging Face upon acceptance.
Figures & tables
| Dataset | Split | Orig. | Comp. | Total |
| ASV19 [ 32 ] | Training | 25,380 | 177,660 | 203,040 |
| ASV19 [ 32 ] | Validation | 24,986 | 174,902 | 199,888 |
| ASV19 [ 32 ] | Evaluation | 71,933 | 503,531 | 575,464 |
| FoR [ 33 ] | Evaluation | 4,634 | 32,438 | 37,072 |
| Wild [ 34 ] | Evaluation | 31,780 | 222,460 | 254,240 |
| Total | 158,713 | 1,110,991 | 1,269,704 | |
| Model | Dataset | Original | MP3/AAC/Opus | BigCodec | DAC | HiggsV2 | Mimi | SNAC | Speech Tokenizer | Wav Tokenizer | Average | |||
| 12K | 24K | 48K | 96K | |||||||||||
| AASIST ∗ [ 35 ] | ASV19 | 0.83 | 13.13 | 2.30 | 1.01 | 0.91 | 8.35 | 1.62 | 4.53 | 3.02 | 12.03 | 5.10 | 18.15 | 7.54 |
| Wild | 43.02 | 51.51 | 45.79 | 43.73 | 43.26 | 34.91 | 39.55 | 34.99 | 36.39 | 32.88 | 42.28 | 47.28 | 38.33 | |
| FoR | 44.13 | 26.80 | 16.03 | 15.49 | 16.12 | 21.34 | 21.06 | 25.83 | 20.22 | 26.37 | 22.96 | 31.1 | 24.13 | |
| XLS-R-SLS ∗ [ 36 ] | ASV19 | 0.24 | 7.33 | 3.59 | 0.72 | 0.24 | 4.80 | 0.52 | 1.93 | 1.14 | 13.30 | 3.21 | 18.59 | 6.21 |
| Wild | 7.46 | 27.43 | 13.32 | 9.91 | 9.22 | 25.43 | 10.17 | 17.44 | 13.51 | 37.03 | 27.81 | 46.62 | 25.43 | |
| Codec Method(s) | Dataset | Original | MP3/AAC/Opus | BigCodec | DAC | HiggsV2 | Mimi | SNAC | Speech Tokenizer | Wav Tokenizer | Average | |||
| 12K | 24K | 48K | 96K | |||||||||||
| Baseline | ASV19 | 0.34 | 7.97 | 5.65 | 0.83 | 0.38 | 5.11 | 0.75 | 2.35 | 1.74 | 13.39 | 4.02 | 21.10 | 6.92 |
| Wild | 8.28 | 25.13 | 10.94 | 8.02 | 7.75 | 23.68 | 8.65 | 15.31 | 12.05 | 36.87 | 24.84 | 43.31 | 23.53 | |
| FoR | 11.35 | 35.78 | 24.47 | 14.48 | 9.58 | 31.77 | 22.49 | 49.50 | 32.41 | 61.31 | 42.58 | 53.45 | 41.93 | |
| BigCodec [ 25 ] | ASV19 | 0.14 | 5.98 | 1.28 | 0.23 | 0.19 | 2.55 | 0.41 | 1.44 | 1.11 | 12.67 | 2.88 | 15.25 | 5.19 |
| Wild | 9.29 | 27.01 | 12.67 | 10.45 | 10.10 | 17.56 | 9.49 | 14.54 | 12.91 | 26.63 | 20.86 | 32.46 | 19.21 | |
| Model | CodecFake+ [ 14 ] | CDD [ 23 ] |
|---|---|---|
| AASIST [ 35 ] | 42.23 | 37.72 |
| XLS-R-SLS [ 36 ] | 24.26 | 31.32 |
| CLAD [ 2 ] | 44.57 | 41.95 |
| RawNet2 [ 21 ] | 44.9 | 40.60 |
| W2W2-AASIST [ 13 ] | 9.9 | 16.89 |
| MMS-300M [ 37 ] | 5.5 | 6.51 |