cs.SDOct 1, 2026

Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR

Authors: Hwayeon Kim, Youngwon Choi, Hyeonyu Kim

Organizations: MAUM AI Inc., Republic of Korea

Abstract

Recent ASR development has placed growing emphasis on generalization across diverse domains and acoustic conditions. Existing approaches typically adapt pretrained ASR models to front-end functions such as wake-up word (WuW) detection through additional training or task-specific modules. In this work, we explore the use of a shared pretrained ASR backbone for WuW detection without gradient-based fine-tuning and examine whether a compact encoder can be extracted using the PCA-based structured pruning approach of SliceGPT. Experiments with Parakeet-TDT-0.6B-v3 and Moonshine-base show that WuW detection performance remains relatively stable when the encoder channel dimension is reduced by 50%. These results suggest that task-relevant compact encoders can be derived from pretrained ASR models without fine-tuning.

Figures & tables

Explore similar work

CardsList
  1. Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

    Sep 23, 2026Rasmus Aagaard, Nicki Skafte DetlefsenEncodersDecode Speedup

  2. Pruning as Regularization: Sensitivity-Aware One-Shot Pruning in ASR

    Nov 11, 2025Julian Irigoyen, Arthur Söhler, Andreas Søeborg KirkedalAutomatic Speech RecognitionRegularization

  3. Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR

    Jul 6, 2026Ho Lam Chung, Yiming Chen, Dau-Cheng Lyu +2End-To-End Automatic Speech Recognition ModelsWhisper-Large-V3