GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection
Organizations: The Hong Kong University of Science and Technology (Guangzhou)
Abstract
Recent advances in AI-based speech synthesis have enabled highly realistic speech, increasing the importance of speech deepfake detection (SDD) in preventing misuse. While mainstream Self-Supervised Learning (SSL)-based detectors achieve strong performance, they suffer from poor generalization to unseen domains and often overlook fine-grained signal artifacts due to a bias towards global semantic consistency. In this paper, we conduct the first detailed empirical and visual analysis to validate these limitations explicitly. Our investigation reveals two critical architectural vulnerabilities: (1) a systemic failure to capture localized spoofing traces, and (2) a severe lack of adaptability to domain-driven shifts in SSL layer importance, rendering static aggregation strategies prone to overfitting. To address these vulnerabilities, we propose the Global-Local Adaptive Detector (GLAD). Specifically, to capture localized forgeries, GLAD employs a Hierarchical Global-Local (HGL) backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details. To counter layer importance shifts in out-of-distribution (OOD) scenarios, we introduce a Hierarchical Adaptive Gating (HAG) mechanism that dynamically recalibrates layer-wise focus in a sample-specific manner. Finally, to address shortcut learning induced by environmental biases, we introduce SaniBoost, a composite data augmentation strategy for robust signal standardization and noise sanitization. Extensive experiments demonstrate that GLAD significantly outperforms state-of-the-art methods, particularly on unseen domain cases.The code will be released upon publication.
Figures & tables
| Method | 2021 LA | 2021 DF | PartialSpoof | ITW | FoR | |
|---|---|---|---|---|---|---|
| EER (%) | min t-DCF | EER (%) | EER (%) | EER (%) | EER (%) | |
| AASIST | 12.82 | 0.5468 | 20.25 | 33.02 | 43.50 | 44.26 |
| W2V2-ST | 7.36 | 0.3769 | 8.13 | 10.96 | 18.67 | 8.98 |
| AMSDF | 2.00 | 0.2408 | 3.82 | 14.07 | 15.45 | 11.44 |
| XLS-53 & LGF | 6.53 | 0.3400 | 4.75 | 9.25 | 26.37 | 21.90 |
| XLS-R & SLS | 4.47 | 0.2907 | 2.30 | 8.28 | 7.38 | 5.15 |
| Method | 21LA | 21DF | ITW |
|---|---|---|---|
| w/o HGL | 3.13 | 1.99 | 9.91 |
| w/o HAG | 2.35 | 2.66 | 9.05 |
| w/o Saniboost | 4.62 | 2.47 | 6.44 |
| GLAD | 1.88 | 1.31 | 5.44 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | 2021 LA EER | 2021 LA min-tDCF | 2021 DF | PartialSpoof | ITW | FoR |
|---|---|---|---|---|---|---|
| AASIST | ||||||
| W2V2-ST | ||||||
| AMSDF | ||||||
| XLSR-LGF | ||||||
| XLSR-SLS | ||||||
| XLSR-MultiConv |
| Method | A07 | A08 | A09 | A10 | A11 | A12 | A13 | A14 | A15 | A16 | A17 | A18 | A19 | EER | min t-DCF |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TSSDN | 1.43 | 0.75 | 0.02 | 1.75 | 0.06 | 0.18 | 0.06 | 0.11 | 2.05 | 1.25 | 6.01 | 1.14 | 1.44 | 1.62 | 0.0474 |
| RawGAT-ST | 0.11 | 0.30 | 0.04 | 0.16 | 0.09 | 0.19 | 0.07 | 0.07 | 0.14 | 0.27 | 0.53 | 0.1 | 0.23 | 1.15 | 0.0373 |
| AASIST | 0.80 | 0.44 | 0.00 | 1.06 | 0.16 | 0.31 | 0.91 | 0.1 | 0.15 | 0.65 | 0.72 | 1.52 | 3.40 | 0.62 | 1.13 |
| LC-Res18 | 0.05 | 0.26 | 0.00 | 0.26 | 0.13 | 0.17 | 0.18 | 0.10 | 0.18 | 0.10 | 2.48 | 0.40 | 0.17 | 0.80 | 0.0210 |
| W2V2-ST | 0.06 | 0.06 | 0.02 | 0.40 | 0.10 | 0.14 | 0.00 | 0.06 | 0.24 | 0.06 | 0.37 | 0.84 | 0.35 | 0.25 | 0.0071 |
| AMSDF | 0.03 | 0.05 | 0.03 | 0.52 | 0.38 | 0.23 | 0.02 | 0.05 | 0.33 | 0.12 | 0.30 | 0.47 | 0.41 | 0.31 | 0.0097 |
| Model / Paper | Parameters / Backbone | 21LA | 21DF | ITW | FoR |
|---|---|---|---|---|---|
| AWaveFormer Wang et al. (2025) | 634M, XLS-R + WavLM | 2.33 | 3.63 | 10.25 | 5.15 |
| Towards Generalisable and Calibrated Audio Deepfake Detection Pascu et al. (2024) | 2.16B, XLS-R-2B | – | – | 7.20 | 6.30 |
| A Robust Audio Deepfake Detection System via Multi-view Feature Yang et al. (2024) | 940M, XLS-R + WavLM + HuBERT | 6.56 | – | – | – |
| Comparative Analysis of ASR Methods for Speech Deepfake Detection Salvi et al. (2024) | 769M, Whisper Medium | – | 12.40 | 32.19 | 12.50 |
| Comparative Analysis of ASR Methods for Speech Deepfake Detection Salvi et al. (2024) | 1.55B, Whisper Large | – | 11.68 | 30.40 | 5.08 |
| ALLM4ADD Gu et al. (2025) | 7.7B | – | – | 26.99 | – |
| Frame AUPRC | IoU@GT | Segment Recall | ||||
|---|---|---|---|---|---|---|
| # segments | XLS-R-SLS | GLAD | XLS-R-SLS | GLAD | XLS-R-SLS | GLAD |
| 1 | 0.1734 | 0.2034 | 0.0822 | 0.0898 | 0.2360 | 0.3202 |
| 2 | 0.3025 | 0.3894 | 0.1636 | 0.2346 | 0.4022 | 0.5858 |
| 3 | 0.2224 | 0.2786 | 0.1135 | 0.1458 | 0.2545 | 0.3737 |
| 4 | 0.2072 | 0.2553 | 0.1050 | 0.1340 | 0.2232 | 0.3406 |
| 5 | 0.1590 | 0.2109 | 0.0801 | 0.1086 | 0.1743 | 0.2770 |
| Dataset | Dynamic | Target-domain Mean | EER [95% CI] |
|---|---|---|---|
| 21LA | 5.3030 | 5.7285 | +0.4255 [0.3268, 0.5223] |
| 21DF | 1.3108 | 1.4709 | +0.1601 [0.1277, 0.2048] |
| ITW | 5.4418 | 5.6195 | +0.1777 [0.0592, 0.2877] |
| Dataset | Layer (%) | Stream (%) | Local (%) |
|---|---|---|---|
| 21LA | 61.2 | 55.9 | 98.6 |
| 21DF | 66.0 | 74.0 | 99.8 |
| ITW | 82.2 | 65.0 | 98.1 |
| Measurement | Median [25th, 75th percentile] |
|---|---|
| Energy change below 500 Hz | dB |
| Energy change, 500–2000 Hz | dB |
| Energy change, 2000–4000 Hz | dB |
| Energy change, 4000–8000 Hz | dB |
| Spectral centroid change | kHz |
| RMS amplitude change | dB |
| Measurement | Original | Forced SaniBoost | Change |
|---|---|---|---|
| Word errors / reference words | 48/3523 | 132/3523 | errors |
| Corpus WER (%) | 1.362 | 3.747 | |
| Corpus CER (%) | 0.611 | 1.957 | |
| 95% CI for the WER increase | |||
| Unchanged utterance-level WER | 431/500 (86.2%) | ||
| Increased utterance-level WER | 64/500 (12.8%) | ||
| Mean Paired Cosine | Linear CKA | |||
|---|---|---|---|---|
| Stage | ||||
| WavLM | 0.8965 | 0.8858 | 0.7522 | 0.6480 |
| XLS-R | 0.6896 | 0.8704 | 0.0877 | 0.6764 |
| Cross-SSL | 0.7325 | 0.8724 | 0.1327 | 0.6597 |
| CNN | 0.9783 | 0.9913 | 0.1794 | 0.1787 |
| Multi-stream fusion | 0.3095 | 0.9096 | 0.0407 | 0.3705 |
| Measurement | ||
| Original inputs classified as bona fide | 498/500 | 499/500 |
| Sanitized inputs classified as bona fide | 278/500 | 500/500 |
| Conditional bona fide retention | 277/498 | 499/499 |
| All predictions unchanged | 278/500 | 499/500 |
| Spoof spoof | 1 | 0 |
| Spoof bona fide | 1 | 1 |
| Measurement | Result |
|---|---|
| Median paired speaker cosine | 0.884 |
| Interquartile interval | |
| Mean paired speaker cosine | 0.860 |
| Dataset | Training condition | Original EER | Conflict EER | EER |
|---|---|---|---|---|
| 21LA | No augmentation | 4.50 | 71.44 | |
| Without ABFP | 2.41 | 11.52 | ||
| Full SaniBoost | 1.81 | 0.02 | ||
| 21DF | No augmentation | 2.41 | 64.59 | |
| Without ABFP | 1.96 | 23.86 | ||
| Full SaniBoost | 1.20 | 0.04 |
| Training condition | SNR (dB) | EER | EER |
| 21LA | |||
| No augmentation | 10 | 63.43 | |
| 5 | 71.44 | ||
| 0 | 74.46 | ||
| Without ABFP | 10 | 10.15 | |
| 5 | 11.52 | ||
| Dataset | Observed EER contrast | 95% CI |
|---|---|---|
| 21LA | ||
| 21DF | ||
| ITW |
| Method / Variant | 21LA | 21DF | ITW |
| w/o L-A Cross-Attention | 3.82 | 12.31 | 12.82 |
| w/o Multi-Granularity Fusion | 2.21 | 2.05 | 8.33 |
| w/o Adaptive S-w Gating | 2.26 | 1.71 | 7.67 |
| w/o Refined L-w Aggregation | 4.70 | 1.96 | 6.14 |
| w/o Global-Local Modulation | 2.37 | 1.66 | 6.19 |
| w/ RawBoost | 3.64 | 1.90 | 7.80 |
| Ratio | DF EER | LA EER/min t-DCF | ITW EER |
|---|---|---|---|
| 0.05 | 1.94 | 1.65 / 0.2314 | 7.96 |
| 0.15 | 1.82 | 2.85 / 0.2550 | 6.47 |
| 0.25 | 1.91 | 5.33 / 0.3299 | 6.51 |
| 0.35 | 1.31 | 1.88 / 0.2359 | 5.44 |
| 0.45 | 2.28 | 2.61 / 0.2573 | 5.63 |
| 0.50 | 1.94 | 3.30 / 0.2719 | 6.97 |
| Cutoff Frequency (Hz) | DF EER | LA EER/min t-DCF | ITW EER |
|---|---|---|---|
| 0 | 1.59 | 3.96 / 0.2936 | 6.42 |
| 1000 | 1.92 | 2.74 / 0.2555 | 6.71 |
| 2000 | 1.31 | 1.88 / 0.2359 | 5.44 |
| 3000 | 2.01 | 3.55 / 0.2835 | 5.95 |
| 4000 | 1.69 | 4.40 / 0.3025 | 5.42 |
| 5000 | 1.87 | 4.09 / 0.2923 | 4.98 |
| Model | Total / Trainable Parameters (M) | Peak Training Memory (GiB) | FLOPs (G) | 21LA EER (%) | 21DF EER (%) | ITW EER (%) |
|---|---|---|---|---|---|---|
| Naive dual-SSL + CNN fusion | 633.89 / 633.89 | 15.58 | 375.01 | |||
| Static SLS + late fusion | 633.90 / 633.90 | 15.76 | 375.02 | |||
| Generic cross-attention + SE | 642.55 / 642.55 | 15.67 | 378.39 | |||
| Parameter-matched large MLP | 662.24 / 662.24 | 15.89 | 375.10 | |||
| MFA-style pooling + late fusion | 646.71 / 646.71 | 16.74 | 377.42 | |||
| GLAD | 662.23 / 662.23 | 16.14 | 381.58 |