Neural Audio Codec for Robust Audio Deepfake Detection
Authors: Jungwoo Kim, Joonyong Park, Junyoung Koh, Jong-Seok Lee
Organizations: Yonsei University, Seoul, Republic of Korea · MAAP Lab, Republic of Korea · The University of Tokyo, Tokyo, Japan · University of Michigan, Ann Arbor, MI, USA
Audio deepfake detectors are typically evaluated on uncompressed audio, although real-world audio often undergoes low-bitrate coding. In this work, we investigate how audio coding affects deepfake detection across codecs, bitrates, and detectors, finding higher errors at lower rates. A mixed-pair protocol isolates codec-induced changes in bona fide and spoof audio, revealing asymmetric, codec-dependent failures: low-rate DAC and EnCodec mainly degrade bona fide detection, whereas X-Codec shows a stronger spoof-side limitation. Motivated by these, we propose a forensic-preserving neural audio codec (FP-NAC), which fine-tunes a pretrained codec using a detector-guided objective while preserving its native hard quantization path and bitrate. On ASVspoof 2019 LA, FP-NAC reduces EER by up to 49.8pp compared with the original DAC at 0.5kbps while maintaining comparable reconstruction quality. Although supervised by only one detector, FP-NAC improves performance across multiple detectors, highlighting forensic transparency as a codec design objective alongside perceptual quality. Our codes are available at https://github.com/kjungwoo03/FP-NAC.
Figures & tables
Figure 1 : Low-bitrate Coding Degrades Audio Deepfake Detection. Full EER (%) on the ASVspoof 2019 LA evaluation set [ 24 ] is reported with its measured bitrate across three different detectors. The gray horizontal lines indicate the EER on uncompressed audio. EER generally increases as bitrate decreases, while FP-NAC substantially reduces the EER compared with other NACs.
Figure 2 : Codec-Dependent Asymmetric Failure Modes. H1 and H2 across bitrates under AASIST [ 9 ] . Acoustic codecs (DAC, EnCodec) show large H1 that shrinks with bitrate, while X-Codec shows a persistently elevated H2 .
Figure 3 : Class-Directed Representation Shifts in AASIST. Detector embeddings at 0.5 kbps are projected onto the clean bona fide–spoof centroid axis. DAC strongly shifts bona fide representations toward spoof, while FP-NAC substantially mitigates this shift.
Figure 4 : Overall Architecture of FP-NAC. The left panel shows frozen DAC backbones with trainable residual adapters conditioned on the active RVQ depth K . The right panel illustrates the adapter architecture, consisting of bottleneck projection, depthwise temporal convolutions, K -conditioned FiLM modulation, and residual addition.
Figure 5 : Multi-Rate Sampling Distribution. Sampling probability over the active RVQ depth K .
Figure 6 : Reconstruction Quality across Bitrates. PESQ-WB [ 8 ] and STOI [ 21 ] on a subset of ASVspoof 2019 LA [ 24 ] . Both reconstruction-only adaptation and FP-NAC improve low-rate reconstruction quality over Plain DAC, with closely matched quality across most operating points.
Configuration
Trainable Params.
Overhead
Plain DAC
74.14M
0%
Encoder only
+0.72M
+0.97%
Decoder only
+0.58M
+0.78%
Both (FP-NAC)
+1.30M
+1.75%
Table 1 : Parameter Overhead of Adapters. Additional trainable parameters relative to the 74.14M-parameter original DAC [ 12 ] .
Figure 7 : Encoder–Decoder Adaptation Ablation at 0.5 kbps. Full EER (%) across AASIST, AASIST-L, and RawNet2 for encoder-only, decoder-only, and joint adaptation (ours).
Department of Engineering Physics Tsinghua University Beijing, China · Tsinghua Shenzhen International Graduate School Tsinghua University Shenzhen, China