Adapting multilingual speech foundation models to low-resource languages remains difficult, especially for languages that are poorly represented during pre-training. While parameter-efficient fine-tuning (PEFT) reduces the cost of adapting large models, conventional approaches such as LoRA rely on generic low-rank parameterizations and do not explicitly use downstream task information to define the adaptation subspace. To investigate whether task-informed PEFT can better support low-resource ASR, we apply Fisher-Whitened Cross-Covariance Analysis (FCCA) to Whisper and Qwen3-ASR, and introduce two complementary extensions: Asymmetric-Coupled FCCA (AC-FCCA), which exploits structured cross-layer sharing, and Adaptive-Rank FCCA (AR-FCCA), which reallocates adaptation capacity across projection matrices under a fixed parameter budget. Under controlled multilingual experiments, we evaluate these approaches on languages that are poorly represented or unsupported during pre-training alongside well-represented languages. Standard FCCA is competitive with, and usually outperforms, trainable-parameter-budget-matched LoRA. AR-FCCA provides the most consistent improvement over standard FCCA across both model architectures, with statistically significant gains in several evaluation settings, while retaining the same number of trainable parameters. These results show that task-informed subspace construction can be effective for low-resource speech adaptation, and that adaptive rank allocation provides a robust way to improve parameter efficiency without increasing model capacity.
Figures & tables
Language
Lang. Token
Train
Val
Test
Asturian
⟨∣es∣⟩
7.5
0.9
2.4
Sorani Kurdish
⟨∣fa∣⟩
10.5
1.2
3.0
Kyrgyz
⟨∣kk∣⟩
9.3
1.3
3.2
Mandarin
⟨∣zh∣⟩
9.7
1.3
3.1
Persian
⟨∣fa∣⟩
10.0
1.5
3.7
Table 1 : FLEURS dataset statistics (hours) for target languages.
Unsupported
Supported
Method
%Para
Ast.
Sor.
Kyr.
Man.
Per.
Vanilla
0
51.11
114.03
90.46
8.88
47.30
FFT
100
15.60
35.53
18.39
9.97
14.20
LoRA-128
9.82
15.96
37.83
19.75
10.10
14.92
LoRA-8
0.61
19.17
44.29
25.63
10.66
16.16
FCCA
0.61
17.07 ∗
44.97
22.25 ∗
10.46
16.20
Table 2 : WER (%) on FLEURS test sets using Whisper medium; CER (%) for Mandarin.
Method
%Para
Ast.
Sor.
Vanilla
0
48.94
105.30
FFT
100
16.06
37.93
LoRA-128
4.50
16.05
40.78
LoRA-6
0.21
19.78
44.99
FCCA
0.20
19.61
44.22
AC-FCCA-H
19.63
44.36
Table 3 : WER (%) on FLEURS test sets using Qwen3-ASR-1.7B.
Figure 1 : Adaptive FCCA rank allocation across Whisper model layers and attention projections for Asturian. Darker cells indicate larger allocated rank.