Many organizations fine-tune publicly available pretrained Automatic Speech Recognition (ASR) models and deploy them in black-box settings, assuming limited access provides protection. We show this assumption is fragile: adversarial perturbations crafted on the public base model transfer effectively to fine-tuned target models, severely degrading performance and posing concerns for safety-critical applications. We propose TransferBreaker, a unified fine-tuning framework that suppresses adversarial transfer by integrating Base Adversarial Fine-Tuning, which restricts adversarial training to base-effective perturbations; Latent Jacobian Regularization, which enforces latent-space invariance by suppressing adversarially sensitive directions; and HybridGrad-AFT, which improves robustness against adaptive attacks by interpolating transferable perturbations from base and target gradients. We theoretically justify all components and evaluate TransferBreaker across three languages and four large ASR models, reducing adversarial WER from 92.6 to 27.8. Our code is publicly available at https://github.com/rohban-lab/TransferBreaker.
Figures & tables
Figure 1 : Adversarial Transferability from Base to Fine-Tuned ASR Models. Fine-tuned (FT) ASR models evaluated across multiple languages and backbones exhibit a sharp increase in Word Error Rate (WER) on adversarial examples optimized against the public base model, relative to clean inputs, revealing a critical transferability vulnerability. In contrast, our proposed framework, TransferBreaker , effectively suppresses this transferability and preserves robustness against base-model-derived attacks. Note: WER can exceed 100% due to insertion/deletion errors; see Appendix J .
Figure 2 : The TransferBreaker Framework. Left: Adversarial perturbation generation, including Base Perturbation ( δbase ) and HybridGrad Perturbation ( δhybrid ), where the latter combines base and target gradients to interpolate transferable adversarial directions. Right: The training loop using clean and adversarial samples with task loss (CTC or Cross-Entropy) and Latent Jacobian Regularization to enforce latent-space invariance and suppress adversarially sensitive directions.
Figure 3 : Motivation for Latent Jacobian Regularization. For Base-AFT target models on Polish, adversarial CER is positively associated with clean-to-adversarial RMS latent distance. Each point corresponds to an evaluation utterance, and the dashed lines show regression fits. This association motivates explicitly constraining adversarial latent displacement during fine-tuning.
Whisper Medium
Wav2Vec2 Large XLSR
MMS 1B
Whisper Large V3 Turbo
Language
Setup
Metric
Base
FT
Std-AFT
Ours
Base
FT
Std-AFT
Ours
Base
FT
Std-AFT
Ours
Base
FT
Std-AFT
Ours
Polish
Clean
WER
15.3
9.6
12.9
12.0
18.8
12.2
17.3
16.8
12.9
8.9
12.4
10.7
13.4
9.6
10.8
11.8
CER
5.9
2.7
4.2
3.8
4.4
3.0
4.5
4.2
3.1
2.1
3.2
2.6
4.3
2.6
3.3
3.7
Base-Adv
WER
118.1
86.7
44.0
24.2
102.0
86.2
33.7
24.8
104.7
84.0
28.7
17.2
128.0
68.4
32.2
20.8
CER
83.2
46.0
22.5
9.7
54.7
46.2
10.3
6.9
63.7
44.3
8.8
4.8
78.8
32.8
14.2
7.8
Portuguese
Clean
WER
22.2
12.5
16.1
15.1
30.0
19.8
26.4
23.9
19.8
14.2
18.2
16.4
16.9
12.6
12.9
12.7
Table 1 : Main results comparing WER/CER under clean and transferred adversarial settings across four ASR backbones. Base model and its variants are trained under three regimes: fine-tuning (FT), standard adversarial fine-tuning (Std-AFT), and TransferBreaker (Ours), using three languages from Common Voice 24. “Base-Adv” denotes adversarial examples crafted on the base model and transferred to the corresponding target models. Lower is better. Note: WER can exceed 100% due to insertion/deletion errors; see Appendix J .
Whisper Medium
Wav2Vec2 Large XLSR
MMS 1B
Language
Setup
Metric
Base-AFT
Base+Std-AFT
Ours
Base-AFT
Base+Std-AFT
Ours
Base-AFT
Base+Std-AFT
Ours
Polish
Clean
WER
11.0
13.1
12.0
15.9
16.7
16.8
10.6
11.0
10.7
CER
3.4
4.2
3.8
3.9
4.2
4.2
2.6
2.6
2.6
Base-Adv
WER
22.4
30.2
24.2
24.0
28.9
24.8
17.2
21.1
17.2
CER
8.3
13.2
9.7
6.9
8.5
6.9
4.9
6.2
4.8
Surrogate-Adv
WER
57.6
46.8
44.1
54.2
44.7
40.7
48.6
38.0
37.2
Table 2 : Comparison of Base-AFT, Base+Standard-AFT, and TransferBreaker under clean settings, base-model adversarial attacks (Base-Adv), and surrogate-based adaptive adversarial attacks (Surrogate-Adv) for Polish and Portuguese across ASR backbones. Lower is better.
Table 6
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Whisper Medium
Wav2Vec2 Large XLSR
MMS 1B
Language
Metric
w/o LJR
w/ LJR
w/o LJR
w/ LJR
w/o LJR
w/ LJR
Polish
Ex[∥Jht(x)δ∥22]
438,244
147,456
5,041
529
30,976
17,424
Portuguese
Ex[∥Jht(x)δ∥22]
338,724
96,721
3,600
625
28,224
12,769
Appendix
Table 5 : Expected latent sensitivity to input perturbations Ex[∥Jht(x)δ∥22] across ASR backbones, with and without LJ-Regularization (LJR). Lower is better.
Figure 4 : Clean and base-model adversarial attack (Base-Adv) WER across different fine-tuning data sizes from the ParlaSpeech Polish dataset using Whisper-Medium and Wav2Vec2-Large-XLSR. Lower is better.
Whisper Medium
Wav2Vec2 Large XLSR
MMS 1B
Language
Setup
Metric
Base-AFT
Base-AFT +LJR
Base-AFT +HybridGrad-AFT
Full
Base-AFT
Base-AFT +LJR
Base-AFT +HybridGrad-AFT
Full
Base-AFT
Base-AFT +LJR
Base-AFT +HybridGrad-AFT
Full
Polish
Clean
WER
11.0
10.8
12.8
12.0
15.9
16.1
16.4
16.8
10.6
10.6
10.5
10.7
CER
3.4
3.3
4.1
3.8
3.9
3.9
4.0
4.2
2.6
2.7
2.6
2.6
Base-Adv
WER
22.4
19.5
28.4
24.2
24.0
22.5
26.5
24.8
17.2
15.6
19.2
17.2
CER
8.3
7.0
12.4
9.7
6.9
6.4
7.8
6.9
4.9
4.4
5.5
4.8
Surrogate-Adv
WER
57.6
56.6
45.8
44.1
54.2
53.2
42.5
40.7
48.6
47.9
38.5
37.2
Appendix
Table 6 : Four-way component ablation on Polish and Portuguese across three ASR backbones. Full denotes the complete TransferBreaker framework, combining Base-AFT, LJR, and HybridGrad-AFT; Base-Adv and Surrogate-Adv denote attacks generated on the public base model and a defense-aware surrogate, respectively. Lower is better.
Whisper Medium
Wav2Vec2 Large XLSR
Language
Setup
w/o
Self
Base (ours)
w/o
Self
Base (ours)
Polish
Clean
12.8/4.1
12.5/3.9
12.0/3.8
16.4/4.0
16.6/4.1
16.8/4.2
Base-Adv
28.4/12.4
26.6/11.2
24.2/9.7
26.5/7.8
25.9/7.5
24.8/6.9
Portuguese
Clean
15.9/5.3
15.5/5.0
15.1/4.8
23.3/6.6
23.6/6.7
23.9/6.7
Base-Adv
26.0/10.9
24.0/9.5
21.7/8.2
34.5/11.5
33.8/11.2
32.7/10.7
Appendix
Table 7 : Frozen-base anchoring versus stop-gradient self-invariance. “w/o” removes LJR, “Self” uses the stop-gradient objective in Eq. 15 , and “Base” uses the frozen-base objective in Eq. 12 . Values are WER/CER; lower is better.
Figure 5 : Performance under base - model adversarial attacks across varying SNR levels, evaluated on Polish using Whisper - Medium and Wav2Vec2 - Large - XLSR. Lower is better.
Whisper Medium
Wav2Vec2 Large XLSR
MMS 1B
Language
Attack
Metric
Ours
Ours
Ours
Polish
PGD1000
WER
22.5
23.8
17.0
CER
9.0
6.6
4.7
C&W
WER
23.5
23.2
16.5
CER
9.4
6.6
4.6
AutoAttack
WER
22.7
22.8
16.3
Appendix
Table 8 : WER/CER (%) of TransferBreaker ( Ours ) under transferred attacks (PGD1000, C&W, AutoAttack, SAGO, M-DI 2 , ILAP, APGD+DI, and A-PGD) on Polish and Portuguese across ASR backbones. Higher is better (stronger attack against our defense).
Model
MI-CWA (Ours)
MI-CWA (Base-AFT)
Surr.-A-PGD (Ours)
Whisper Medium
46.4/21.9
60.2/30.8
44.1/20.5
Wav2Vec2 Large XLSR
42.8/16.3
56.8/24.8
40.7/15.3
MMS 1B
39.1/14.2
51.4/21.0
37.2/13.3
Appendix
Table 9 : WER/CER (%) under the ensemble-based adaptive attack MI-CWA, compared to the single-surrogate adaptive A-PGD, on Polish. Lower is better.
Whisper Medium (Ours)
Wav2Vec2 Large XLSR (Ours)
Language
Setup
Metric
λ1=2
λ1=4
λ1=6
λ1=2
λ1=4
λ1=6
Polish
Clean
WER
11.5
12.0
13.2
16.1
16.8
17.4
CER
3.6
3.8
4.3
4.0
4.2
4.4
Base-Adv
WER
25.9
24.2
24.8
25.6
24.8
24.9
CER
10.5
9.7
10
7.2
6.9
6.9
Appendix
Table 10 : Effect of the adversarial loss coefficient λ1 on clean accuracy and adversarial transferability. Results are reported for Polish using Whisper-Medium and Wav2Vec2-Large-XLSR backbones. Lower is better.
Whisper Medium (Ours)
Wav2Vec2 Large XLSR (Ours)
Language
Setup
Metric
λ2=1
λ2=2.5
λ2=5
λ2=1
λ2=2.5
λ2=5
Polish
Clean
WER
12.4
12.0
12.7
16.5
16.8
17.0
CER
3.9
3.8
4.0
4.1
4.2
4.3
Base-Adv
WER
25.5
24.2
24.0
25.2
24.8
24.7
CER
10.3
9.7
9.7
7.1
6.9
6.9
Appendix
Table 11 : Effect of the LJ-Regularization coefficient λ2 on clean accuracy and adversarial transferability. Results are reported for Polish using Whisper-Medium and Wav2Vec2-Large-XLSR backbones. Lower is better.
Setting
Clean WER/CER
Base-Adv WER/CER
Time (Hours)
Bernoulli Mixing
12.0 / 3.8
24.2 / 9.7
20.8
Always Both
12.2 / 3.8
23.8 / 9.5
40.5
Appendix
Table 12 : Ablation study comparing Bernoulli mixing and always applying both perturbations on Whisper - Medium for Polish under clean and transferred adversarial settings ( ≈ 30 hours of audio, batch size 4, 4 epochs).
Whisper Medium (Ours)
Wav2Vec2 Large XLSR (Ours)
Language
Setup
Metric
γ=0.2
γ=0.5
γ=1.0
γ=0.2
γ=0.5
γ=1.0
Polish
Clean
WER
11.9
12.0
12.3
16.2
16.8
16.7
CER
3.7
3.8
4.0
4.0
4.2
4.2
Base-Adv
WER
23.5
24.2
25.4
24.3
24.8
25.0
CER
8.9
9.7
10.1
6.8
6.9
7.0
Surrogate-Adv
WER
48.5
44.1
45.2
45.1
40.7
41.1
Appendix
Table 13 : Effect of the HybridGrad ratio γ on clean accuracy and adversarial transferability. Results are reported for Polish using Whisper-Medium and Wav2Vec2-Large-XLSR backbones. Lower is better.
Whisper Medium
Wav2Vec2 Large XLSR
MMS 1B
Language
Setup
Metric
Base
Base-AFT
Base
Base-AFT
Base
Base-AFT
Polish
Adv (HybridGrad)
WER
70.0
63.3
69.7
57.6
71.9
50.3
CER
40.4
28.9
32.1
26.5
34.2
20.8
Portuguese
Adv (HybridGrad)
WER
75.6
62.5
85.2
64.4
84.5
62.7
CER
35.9
29.9
42.4
30.6
40.7
28.8
Appendix
Table 14 : WER/CER of base and Base-AFT models under the HybridGrad adaptive attack on Polish and Portuguese, evaluated across ASR backbones. Higher is better.
Whisper Medium
Language
Setup
Metric
Uniform
Reverse
Ours
Polish
Clean
WER
12.7
12.5
12.0
CER
4.0
4.0
3.8
Base-Adv
WER
27.2
26.4
24.2
CER
11.9
11.5
9.7
Surrogate-Adv
WER
49.0
44.8
44.1
Appendix
Table 15 : Comparison of different HybridGrad scheduling strategies evaluated under clean settings, base-model adversarial attacks (Base-Adv), and surrogate-based adaptive adversarial attacks (Surrogate-Adv) on Whisper-Medium. Lower is better.
FT
Ours
Model
Clean
MUSAN
AMR-NB
Clean
MUSAN
AMR-NB
Whisper-Medium
9.6
12.9
17.8
12.0
13.2
15.6
Wav2Vec2 Large XLSR
12.2
22.6
29.1
16.8
20.7
25.1
MMS-1B
8.9
15.6
20.1
10.7
14.1
17.2
Appendix
Table 16 : WER (%) under common non-adversarial acoustic corruptions on the Polish Common Voice test set. MUSAN-derived babble noise is mixed at 15 dB SNR, and AMR-NB compression uses the 6.7 kbps mode. Lower is better.
Whisper Medium
Wav2Vec2 Large XLSR
Metric
FT
Std-AFT
Ours
FT
Std-AFT
Ours
Peak VRAM (GB)
17.2
18.9
21.7
18.1
18.4
19.3
Time / Epoch (Hours)
0.42
2.91
5.20
0.14
1.00
1.80
Appendix
Table 17 : Training cost comparison between normal fine-tuning (FT), standard adversarial fine-tuning (Std-AFT), and TransferBreaker using Whisper-Medium and Wav2Vec2-Large-XLSR. Experiments are conducted on a single RTX-4090 (24 GB) using ≈ 30 hours of Polish audio.
Target Defended By
Language
Model
Base-AFT
Std-AFT
TransferBreaker
Polish
Whisper-Medium
65.8
57.0
57.5
Wav2Vec2
61.9
54.1
54.5
Portuguese
Whisper-Medium
64.7
57.8
58.3
Wav2Vec2
69.8
61.6
61.0
Appendix
Table 18 : WER (%) under a white-box attack generated directly on each fine-tuned target model trained using Base-AFT, Std-AFT, and TransferBreaker. Lower indicates stronger robustness.
Polish
Portuguese
Knowledge
Surrogate Model
Whisper-Medium
Wav2Vec2
Whisper-Medium
Wav2Vec2
Base model only
None (base model)
24.2
24.8
21.7
32.7
Full white-box
TransferBreaker target (ceiling)
57.5
54.5
58.3
61.0
Full white-box
Std-AFT target (reference)
57.0
54.1
57.8
61.6
Defense procedure
MLS surrogate
44.1
40.7
44.7
48.9
Defense procedure + target data
Same-data surrogate (seed only)
48.6
45.2
49.1
53.4
Appendix
Table 19 : Attack strength against TransferBreaker across the surrogate-fidelity spectrum, from a base-model attack to a full white-box attack on the target. Values are WER (%); higher indicates a stronger attack. Same-data surrogate (seed only) is shown for reference.
Whisper Medium
Wav2Vec2 Large XLSR
Language
Setup
Metric
Std-AFT(4e)
Std-AFT(8e)
Base-AFT(4e)
Std-AFT(4e)
Std-AFT(8e)
Base-AFT(4e)
Polish
Clean
WER
12.9
12.5
11.0
17.3
17.0
15.9
CER
4.2
4.1
3.4
4.5
4.4
3.9
Base-Adv
WER
44.0
41.5
22.4
33.7
31.4
24.0
CER
22.5
21.3
8.3
10.3
9.8
6.9
Appendix
Table 20 : Effect of increasing training epochs for Standard - AFT on Whisper - Medium and Wav2Vec2 - Large - XLSR for Polish under clean and transferred adversarial settings. Base-AFT (4 epochs) is shown for comparison. Lower is better.
Whisper Medium
Language
Setup
Metric
Std-AFT
CleanUNet+FT
Purifier+FT
Ours
Polish
Clean
WER
12.9
10.9
23.1
12.0
CER
4.2
3.5
11.2
3.8
Base-Adv
WER
44.0
64.2
51.8
24.2
CER
22.5
33.8
28.3
9.7
Portuguese
Clean
WER
16.1
14.2
29.4
15.1
Appendix
Table 21 : Comparison of diffusion purification with fine-tuning (Purifier+FT) and CleanUNet denoising with fine-tuning (CleanUNet+FT) against standard-AFT (Std-AFT) and our method on Whisper-Medium for Polish and Portuguese under clean and transferred adversarial settings. Lower is better.
Whisper Medium
Language
Setup
Metric
Base
FT
Ours
Finnish
Clean
WER
19.8
15.8
17.6
CER
6.8
5.2
5.6
Base-Adv
WER
132.9
96.3
26.0
CER
82.8
51.4
10.1
Vietnamese
Clean
WER
27.4
20.7
22.1
Appendix
Table 22 : Evaluation of TransferBreaker on extremely low-resource languages (Finnish and Vietnamese) using Whisper-Medium on Common Voice 24 under clean and transferred adversarial settings. Lower is better.