Frozen speech foundation models (SFMs) make downstream adaptation efficient: the backbone can stay fixed while a task learns layer fusion and a lightweight classifier. Full adversarial fine-tuning is a standard route to robustness, but generating adversarial examples and updating the backbone for every task sacrifices that efficiency. We ask whether robustness can instead be learned before future tasks are known. For a frozen backbone and linear classifier, robustness can be understood through the interaction between representation stability and decision-boundary margin. This leads directly to our design: we stabilize representations across the hidden layers, rather than only the final layer, while preserving clean representations; after clean adaptation selects the layer mixture, we keep it fixed and enlarge only the classifier margin, without downstream adversarial examples. We evaluate Wav2Vec2, HuBERT, and WavLM Large on four tasks under adaptive 30 dB attacks. Across 12 backbone-task pairs, hierarchical robustification improves robust accuracy by 46.4 pp, while margin refinement adds 4.0 pp for 1.1 pp of clean accuracy. Code and configurations are available at https://github.com/arefmousavi/hierarchical-robust-sfm.
Figures & tables
Figure 1: Overview of the proposed framework. Stage 1 robustifies representations across hidden layers before downstream tasks are known, while preserving clean representations. Clean adaptation then learns task-specific convex layer fusion and a linear classifier with the robust backbone frozen. Stage 2 fixes the backbone and fusion and refines only the classifier using normalized geometric margins.
Clean / robust accuracy (%) ↑
ΔRα
ΔM10
ΔCA
ΔRA
Method
KS
IC
SID
ER
(%) ↓
( ×10−3 ) ↑
(pp) ↑
(pp) ↑
Wav2Vec2 Large (primary backbone)
Standard + CE
95.2 ± 0.56 / 5.5 ± 0.96
94.1 ± 0.47 / 2.7 ± 0.52
87.6 ± 0.91 / 0.5 ± 0.09
67.5 ± 0.52 / 0.8 ± 0.18
0.0
0.0
0.0
0.0
FARE-style + CE
91.7 ± 0.19 / 39.7 ± 0.71
90.5 ± 0.56 / 36.8 ± 0.63
83.9 ± 0.48 / 16.0 ± 0.99
64.9 ± 1.48 / 9.7 ± 0.95
-38.3
0.5
-3.4
23.2
Terminal-layer + CE
92.5 ± 0.54 / 38.2 ± 0.75
91.2 ± 0.17 / 36.1 ± 1.36
84.1 ± 0.92 / 15.6 ± 0.58
64.8 ± 0.62 / 9.8 ± 1.43
-37.5
2.1
-3.0
22.6
Stage 1 + CE
93.7 ± 0.56 / 62.2 ± 0.47
92.8 ± 0.58 / 61.0 ± 0.45
84.4 ± 0.47 / 34.8 ± 0.64
65.2 ± 0.73 / 26.5 ± 1.27
-81.8
8.6
-2.1
43.8
Table 1: Robustness transfer across three backbones and four tasks. Task entries are clean/robust accuracy (CA/RA, %), reported as mean ± SD over three downstream adaptation seeds. The four overbarred columns report changes relative to Standard+CE, averaged across tasks within each backbone; Δ CA and Δ RA are in percentage points, while ΔRα and ΔM10 are defined in Eq. ( 13 ). † indicates use of downstream adversarial examples; Full AT is a task-specific oracle outside the frozen-backbone setting.
(a) Stage 1 variants (+CE)
ΔRα
ΔCA
ΔRA
Variant
(%) ↓
(pp) ↑
(pp) ↑
Terminal layer
-37.5
-3.0
22.6
Mean all-layer
-72.9
-2.2
39.3
STAGE 1, no Lpres.
-93.2
-20.8
13.5
Stage 1
-81.8
-2.1
43.8
Table 2: Primary-backbone mechanism and validity suite. (a) Stage 1 variants with clean CE adaptation and without Stage 2 . (b) CE, hard-margin, and soft-margin adaptation using the Stage 1 backbone. (c) Attack-wise and per-example-union RA for our full Stage 1 + Stage 2 method. All values are averaged over three downstream adaptation seeds; overbars then denote four-task means of changes relative to Standard+CE.
Figure 2: Primary-backbone diagnostics on the IC task. (a) Layer-wise representation drift under terminal-layer, mean all-layer, and Stage 1 stabilization. (b) Robust accuracy across 24–36 dB SNR budgets.
Many organizations fine-tune publicly available pretrained Automatic Speech Recognition (ASR) models and deploy them in black-box settings, assuming limited access provides protection. We show this assumption is fragile: adversarial perturbations crafted on the public base model transfer effectively to fine-tuned target models, severely degrading performance and posing concerns for safety-critical applications. We propose TransferBreaker, a unified fine-tuning framework that suppresses adversarial transfer by integrating Base Adversarial Fine-Tuning, which restricts adversarial training to base-effective perturbations; Latent Jacobian Regularization, which enforces latent-space invariance by suppressing adversarially sensitive directions; and HybridGrad-AFT, which improves robustness against adaptive attacks by interpolating transferable perturbations from base and target gradients. We theoretically justify all components and evaluate TransferBreaker across three languages and four large ASR models, reducing adversarial WER from 92.6 to 27.8. Our code is publicly available at https://github.com/rohban-lab/TransferBreaker.
Mojtaba Nafez, Aref Mousavi, Mohammad Ebrahim Mahdavi +3
Department of Computer Engineering Sharif University of Technology · Department of Computer Engineering University of Isfahan
Large speech foundation models have shown strong potential for speech deepfake detection, but direct fine-tuning is limited by a mismatch between self-supervised pre-training objectives and spoof-specific artifacts. To address this, we propose a mix-frame post-training strategy to create localized spoof-oriented perturbations and use frame-level supervision to encourage the SSL model to learn local inconsistencies that are critical for robust spoof detection. On ASVspoof5, we achieve state-of-the-art EER 4.50% for a single model without data augmentation. On ASVspoof2021 LA/DF, it further achieves only 0.16% absolute EER gap between LA and DF, indicating strong and balanced robustness across distinct distortion conditions. These results show that supervised post-training provides an effective and practical way to adapt speech foundation models for robust deepfake detection.
Zihan Pan, Sailor Hardik, Jinyang Wu
Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (A*STAR), 1 Fusionopolis Way, 138632, Singapore
Supervised fine-tuning (SFT) is widely used to adapt self-supervised speech representations to downstream classification tasks. Small gains observed under a single pretrained checkpoint are often interpreted as method-level improvements, i.e., a higher attainable performance ceiling. We show that such conclusions are not always reliable because SFT outcomes depend strongly on the specific pretrained instance. We conduct a systematic study on 3 SUPERB classification tasks, evaluating 8 SFT variants across 9 pretrained checkpoints from wav2vec~2.0, HuBERT, and WavLM, with multi-seed repetitions on representative base-scale models. We find that the identity of the statistically indistinguishable top-group SFT recipe is often checkpoint-dependent, with limited transferability across pretrained instances. These findings suggest that many reported downstream gains reflect instance and seed dependent elicitation match, rather than universally improving the attainable performance ceiling.
Wangjin Zhou, Yizhou Zhang, Yichi Wang +1
Graduate School of Informatics, Kyoto University, Kyoto, Japan