cs.SDOct 5, 2026

Exposing and Mitigating Neural Codec Vulnerabilities in Audio Deepfake Detection

Authors: Abdullah, Awais Khan, Khalid Mahmood Malik

Organizations: College of Innovation and Technology University of Michigan-Flint Michigan, USA

Abstract

Existing audio deepfake detection (ADD) datasets and detectors are primarily built for vocoder-based synthesis, evaluated against traditional post-hoc perturbations such as MP3/AAC compression or additive noise, applied independently of generation. However, recent speech synthesizers, particularly ALM-based systems, use neural audio codecs both for compression and as the resynthesis reconstructing waveforms from generated tokens, producing artifacts distinct from post-hoc compression. Neural codecs thus play a dual role: some are designed for pure compression under low-bandwidth communication, while others serve as resynthesis components. Despite this dual role, robustness to codec-based compression, unlike post-hoc compression, remains largely unexplored. We expose this gap, showing that state-of-the-art (SOTA) ADD models degrade drastically on codec-compressed speech; in particular, systems trained on Codec Resynthesized data as a proxy for codec-based generation prove most vulnerable, with legitimately compressed bona fide speech often misclassified as fake. To investigate this, we construct the Audio Neural Codec-Spoof dataset by applying seven neural codec algorithms to existing ADD benchmarks, isolating codec-induced resynthesis artifacts as a controlled proxy for codec-based generation. As baseline mitigation, we propose PCL-NET (Pairwise Consistency Learned Network), fine-tuning a pretrained XLS-R (300M) encoder with a pairwise consistency objective that minimizes the representation distance between an utterance's uncompressed and codec-compressed versions, disentangling codec artifacts from the real-versus-fake decision. As a result, PCL-NET reduces average EER under neural codec compression from 28.67% to 12.77%, while preserving competitive CoSG-based deepfake detection performance. We will also make the dataset publicly available on Hugging Face upon acceptance.

Figures & tables

Explore similar work

CardsList
  1. Neural Audio Codec for Robust Audio Deepfake Detection

    Sep 30, 2026Jungwoo Kim, Joonyong Park, Junyoung Koh +1Neural Audio CodecsAffectcodec

  2. What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection

    Sep 27, 2026Jiajun Xu, Menglu Li, Xiao-Ping ZhangAudio Deepfake Detection

  3. Mitigating Proxy-to-Wild Domain Gap in Deepfake Speech

    Jun 5, 2026Xuanjun Chen, Yun-Shing Wu, Wei-Chung Lu +4