Recent ASR development has placed growing emphasis on generalization across diverse domains and acoustic conditions. Existing approaches typically adapt pretrained ASR models to front-end functions such as wake-up word (WuW) detection through additional training or task-specific modules. In this work, we explore the use of a shared pretrained ASR backbone for WuW detection without gradient-based fine-tuning and examine whether a compact encoder can be extracted using the PCA-based structured pruning approach of SliceGPT. Experiments with Parakeet-TDT-0.6B-v3 and Moonshine-base show that WuW detection performance remains relatively stable when the encoder channel dimension is reduced by 50%. These results suggest that task-relevant compact encoders can be derived from pretrained ASR models without fine-tuning.
Figures & tables
Figure 1 : Concept of elastic speech inference. A pretrained ASR encoder is reconfigured by selectively using task-relevant subsets of its internal representations. The activation patterns are schematic and do not correspond to specific channel indices or exact subnetwork configurations.
Figure 2 : Results on the Google Speech Commands v2 test set across channel pruning ratios: (a) EER and (b) Encoder parameter reduction ratio. Dashed lines in (a) indicate the EERs of the original encoders.
Model
Method
EER (%) ↓
FRR (%) ↓
FAR (%) ↓
Parakeet-TDT
PCA
6.87
13.81
0.81
Random
50.08 ± 0.06
99.32 ± 0.24
0.69 ± 0.17
Moonshine
PCA
10.80
22.12
0.86
Random
42.03 ± 1.14
97.46 ± 0.37
0.80 ± 0.09
Table 1 : Comparison of PCA-based reduction and random channel selection at 50% pruning. Mean and std values are computed over 3 random seeds.
Table 2 : Effect of LibriSpeech (LS) and GSC v2 as calibration sources.
Figure 3 : Subspace overlap between PCA subspaces derived from LibriSpeech and GSC v2 across encoder layers at 50% pruning. The solid line shows the mean across GSC v2 keywords, and the shaded region indicates the min–max range.
Configuration
Encoder Params(M)
FA/h ↓
FRR(%) ↓
RTF ↓
Original
20.15
0.47
4.51
0.2134
PCA-50%
12.66
0.26
6.72
0.1718
Table 3 : Streaming performance of Moonshine-base.