Learning as Deepfakes Evolve: RF-Prompt for Continual Audio Deepfake Detection
Organizations: The Hong Kong Polytechnic University, Hong Kong SAR, China · Communication University of China, Beijing, China · Beijing Institute of Technology, Beijing, China
Abstract
Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech. Existing dataset-incremental evaluation changes both real-speech domains and deepfake mechanisms, making their effects difficult to distinguish. We construct five task organizations over identical training, development, and evaluation pools to study these factors under a controlled sample budget. Our proposed Real-Anchored Mechanism-Incremental (RAMI) protocol reflects the practical setting in which available real speech provides a recurring mixed-domain reference while new deepfake mechanisms arrive incrementally. We further propose RF-Prompt, an asymmetric continual prompt-learning method that preserves reusable real-speech knowledge through a shared real prompt and expands mechanism-specific knowledge through inherited fake experts with orthogonal residuals. Input-adaptive soft fusion combines the accumulated experts into a fixed number of injected tokens without requiring task identity at inference. On RAMI, RF-Prompt achieves 10.110% average EER and 10.370% pooled EER, outperforming all evaluated continual-learning baselines. Across the five controlled protocols, RAMI yields the lowest common-average and pooled EER. Component ablations, limited-data experiments, and cross-backbone evaluations further validate the proposed design.
Figures & tables
| Method | Category | Venue | M1 | M2 | M3 | M4 | Avg EER | Pool EER | AF |
|---|---|---|---|---|---|---|---|---|---|
| Sequential | Traditional CL | – | 14.50 | 9.12 | 18.78 | 10.12 | 13.130 | 13.445 | 6.433 |
| EWC ( Kirkpatrick and others, 2017 ) | General CL | PNAS’17 | 13.40 | 8.30 | 17.88 | 9.44 | 12.255 | 12.805 | 5.827 |
| OWM ( Zeng et al., 2019 ) | General CL | Nat. Mach. Intell.’19 | 13.80 | 8.68 | 17.64 | 9.90 | 12.505 | 13.080 | 5.993 |
| RAWM ( Zhang et al., 2023 ) | ADD-specific CL | ICML’23 | 10.22 | 8.02 | 14.96 | 11.98 | 11.295 | 11.730 | 5.753 |
| RWM ( Zhang et al., 2024c ) | ADD-specific CL | AAAI’24 | 13.40 | 8.32 | 17.46 | 9.88 | 12.265 | 12.845 | 5.700 |
| RegO ( Chen and others, 2025 ) | ADD-specific CL | AAAI’25 | 14.46 | 8.22 | 17.72 | 9.64 | 12.510 | 13.000 | 6.060 |
| Protocol | real arrival | fake organization | Common Avg EER | Pool EER |
|---|---|---|---|---|
| 1 | Dataset-wise | Dataset-wise | 16.400 | 16.430 |
| 2 | Dataset-wise | Mechanism-wise | 15.675 | 15.900 |
| 3 | Source-support-matched | Mechanism-wise | 13.230 | 13.110 |
| 4 | Four-domain mixture | Dataset-wise | 11.035 | 11.330 |
| 5 (RAMI) | Four-domain mixture | Mechanism-wise | 10.110 | 10.370 |
| Configuration | Avg EER | Pool EER | AF |
|---|---|---|---|
| RF-Prompt (full) | 10.110 | 10.370 | 4.560 |
| w/o real cosine-anchoring loss | 10.755 | 11.365 | 5.293 |
| w/o residual-orthogonality loss | 12.255 | 12.805 | 7.320 |
| w/o adaptive fusion (uniform mean) | 11.525 | 12.480 | 5.480 |
| w/o orthogonal initialization projection | 11.085 | 11.430 | 5.867 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Protocol | Task 1 | Task 2 | Task 3 | Task 4 |
|---|---|---|---|---|
| 1 | 4,900/4,505/9,832 | 3,200/3,535/8,525 | 4,800/4,800/10,000 | 6,300/6,360/11,643 |
| 2 | 4,800/4,800/10,000 | 4,800/4,800/10,000 | 4,800/4,800/10,000 | 4,800/4,800/10,000 |
| 3 | 4,704/5,526/10,600 | 5,819/5,019/10,380 | 4,800/4,800/10,000 | 3,877/3,855/9,020 |
| 4 | 4,900/4,505/9,832 | 3,200/3,535/8,525 | 4,800/4,800/10,000 | 6,300/6,360/11,643 |
| 5 (RAMI) | 4,800/4,800/10,000 | 4,800/4,800/10,000 | 4,800/4,800/10,000 | 4,800/4,800/10,000 |
| Mechanism | Train | Dev | Eval |
|---|---|---|---|
| M1 | ASV2019: 2,400 | ASV2019: 2,000; ASV5: 400 | ASV2019: 3,890; ASV5: 1,110 |
| M2 | ASV2019: 100; ASV5: 800; AT-ADD: 1,500 | ASV2019: 105; ASV5: 735; AT-ADD: 1,560 | ASV2019: 942; ASV5: 2,030; AT-ADD: 2,028 |
| M3 | CodecFake: 2,400 | CodecFake: 2,400 | CodecFake: 5,000 |
| M4 | AT-ADD: 2,400 | AT-ADD: 2,400 | ASV5: 385; AT-ADD: 4,615 |
| Source | Mechanism | Generator IDs in assignment catalog |
|---|---|---|
| ASV2019 | M1 | A02, A03, A04, A05, A06, A11, A13, A14, A16, A17, A18, A19 |
| ASV2019 | M2 | A01, A07, A08, A09, A10, A12, A15 |
| ASV5 | M1 | A12, A19, A20 |
| ASV5 | M2 | A01, A02, A03, A04, A05, A06, A07, A08, A09, A10, A11, A13, A14, A15, A16, A17, A18, A21, A22, A23, A24, A25, A26, A27, A28, A30, A31, A32 |
| ASV5 | M4 | A29 |
| AT-ADD | M2 | BigVGAN, DiffGANTTS, DiffSpeech_fastdiff, E2TTS, F5TTS, FastDiff, FastPitch_fastdiff, FastSpeech2_fastdiff, GlowTTS, GradTTS, HiFiGAN, Kokoro, MBMelGAN, MelGAN, MeloTTS, OpenVoice2, ParallelWaveGAN, PortaSpeech_normal_fastdiff, SeedVC1, StarGANv2VC, StyleMelGAN, StyleSpeech, StyleTTS2, Tacotron2, VITS, WaveGlow, WaveNet, prodiff_teacher_fastdiff |
| Setting | Value |
|---|---|
| Sample rate / input length | 16 kHz / 64,600 samples |
| Tasks / epochs per task | 4 / 50 |
| Batch size / random seed | 32 / 2026 |
| Optimizer | Adam |
| Adam betas / epsilon | (0.9, 0.999) / 1e-8 |
| Weight decay | 5e-4 |
| Method | After Task 1 | After Task 2 | After Task 3 | After Task 4 | Final Avg | Final Pool |
|---|---|---|---|---|---|---|
| Sequential | 4.66 | 8.00/9.38 | 12.84/12.22/9.32 | 14.50/9.12/18.78/10.12 | 13.130 | 13.445 |
| EWC | 4.66 | 8.28/9.68 | 13.58/12.34/9.14 | 13.40/8.30/17.88/9.44 | 12.255 | 12.805 |
| OWM | 4.66 | 9.34/9.80 | 14.08/12.50/8.80 | 13.80/8.68/17.64/9.90 | 12.505 | 13.080 |
| RAWM | 4.66 | 6.18/5.10 | 8.54/6.62/6.18 | 10.22/8.02/14.96/11.98 | 11.295 | 11.730 |
| RWM | 4.66 | 9.24/10.08 | 13.80/12.38/9.10 | 13.40/8.32/17.46/9.88 | 12.265 | 12.845 |
| RegO | 4.66 | 8.82/10.16 | 14.92/12.88/9.34 | 14.46/8.22/17.72/9.64 | 12.510 | 13.000 |
| real protection | fake construction | Avg | Pool | AF |
|---|---|---|---|---|
| None | None | 12.600 | 12.020 | 5.393 |
| None | Inherited residual | 10.755 | 11.365 | 5.293 |
| Shared Prompt Distillation (SPD) | Inherited residual | 10.910 | 11.180 | 5.467 |
| Prompt cosine | Full-Prompt orthogonality | 11.795 | 12.355 | 6.140 |
| Prompt cosine | Inherited residual | 10.110 | 10.370 | 4.560 |
| Unique fake/task | Method | Avg EER | Pool EER | AF |
|---|---|---|---|---|
| 100 | Sequential | 17.325 | 17.750 | 6.887 |
| EWC | 17.705 | 18.080 | 6.707 | |
| RWM | 17.075 | 17.645 | 6.067 | |
| RegO | 17.765 | 18.210 | 6.040 | |
| RF-Prompt | 12.590 | 12.375 | 3.907 | |
| 500 | Sequential | 15.220 | 15.715 | 7.367 |