cs.SDOct 4, 2026

Task-Aware Joint Pruning and Distillation for Efficient Audio Deepfake Detection

Authors: Miao He, Peng Cheng, Zhongjie Ba, Qing Wen, Li Lu, Xin Yang, Kui Ren

Organizations: The State Key Laboratory of Blockchain and Data Security, Zhejiang University, China · Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, China · Qingdao Institute of Software College of Computer Science and Technology, China University of Petroleum (East China), China · Shanghai Institute for Advanced Study, Zhejiang University, China

Abstract

Advances in speech synthesis have made deepfake speeches increasingly convincing, posing growing threats to security. While self-supervised learning (SSL) based detectors achieve state-of-the-art performance, their computational demands (typically 300M+ parameters) prevent deployment on resource-constrained devices. Existing compression methods, designed mainly for content-centric tasks, struggle to maintain competitive performance when directly adapted to deepfake detection. We propose a Task-Aware Joint Pruning and Distillation framework that combines cross-domain knowledge distillation with movement-guided structured pruning to transfer forgery-discriminative knowledge and preserve critical structures under aggressive compression. Our framework reduces the model to 31.9M parameters with 6.3×\times FLOPs reduction, with an average performance drop of only 1.30% across multiple datasets compared to the uncompressed baseline, demonstrating strong potential for on-device deployment.

Figures & tables

Explore similar work

CardsList
  1. FlowFake: Liquid Networks for Audio Deepfake Detection

    Jun 17, 2026Shivaay Dhondiyal, Divyansh Sharma, Dinesh Kumar VishwakarmaNeural AudioSpeaker

  2. Alethia: A Foundational Encoder for Voice Deepfakes

    Apr 30, 2026Yi Zhu, Brahmi Dwivedi, Jayaram Raghuram +1Deepfake DetectionOpencode

  3. Probing-Guided Layer Selection from Self-Supervised Speech Models for Generalizable Audio Deepfake Detection

    Jun 29, 2026Marjan Beheshti, Majid Rostami, Bo ChenSelf-Supervised Speech ModelsNeural Audio