Voice-based interfaces rely on a wake-up word mechanism to initiate communication with devices. However, achieving a robust, energy-efficient, and fast detection remains a challenge. This paper addresses these real production needs by enhancing data with temporal alignments and using detection based on two phases with multi-resolution. It employs two models: a lightweight on-device model for real-time processing of the audio stream and a verification model on the server-side, which is an ensemble of heterogeneous architectures that refine detection. This scheme allows the optimization of two operating points. To protect privacy, audio features are sent to the cloud instead of raw audio. The study investigated different parametric configurations for feature extraction to select one for on-device detection and another for the verification model. Furthermore, thirteen different audio classifiers were compared in terms of performance and inference time. The proposed ensemble outperforms our stronger classifier in every noise condition.
Figures & tables
Figure 1: Attention-based ensemble architecture. Each front-end is the frozen initial stack of a pre-trained classifier and extracts features at its own temporal resolution ri . Projections align the embeddings to a common temporal length tbins′ , which self-attention then fuses. Front-end 1 is the always-on first-stage model, so its states are reused at no additional cost. Generalizes to N front-ends, shown dashed; we use N=2 .
Model
Params.
Oper.
Size (MB)
cnn-fat2019
5.2M
1196.8M
43.6
cnn-trad-pool2
69.5k
10.1M
0.9
gru-att
643.9k
21.9M
5.4
gru-max
145.6k
21.5M
0.8
sgru
145.6k
144.4k
0.8
resnet15
237.4k
456.6M
10.6
Table 1: Parameters, number of operations (multiplications and additions) and size of WuW detection models.
Figure 2: Impact of temporal resolution and number of MFCCs on the WuW F1-score (positive class) by SNR range.
Features
Size
Inference time (ms)
13, w=100, h=50
(29, 13)
25.08
13, w=100, h=20
(71, 13)
51.62
13, w=30, h=10
(148, 13)
73.43
13, w=20, h=10
(149, 13)
82.75
40, w=30, h=10
(148, 40)
83.97
Table 2: MFCC size and inference time on the Pixel XL Android device. The features column shows coefficients, window size ( w ) and hop size ( h ) in milliseconds.
SNR[dB]
device-sgru
resnet15
ensemble
WuW f1-score
95% CI
WuW f1-score
95% CI
WuW f1-score
95% CI
[15, 20]
0.988
0.977, 0.996
0.991
0.981, 0.998
0.989
0.976, 0.996
[10, 15]
0.984
0.969, 0.994
0.993
0.984, 0.999
0.988
0.979, 0.995
[5, 10]
0.978
0.961, 0.990
0.991
0.978, 0.999
0.993
0.989, 0.997
[0, 5]
0.956
0.934, 0.973
0.982
0.967, 0.992
0.983
0.972, 0.989
[-5, 0]
0.888
0.841, 0.923
0.944
0.918, 0.965
0.946
0.929, 0.968
Table 3: WuW F1-score (positive class) and the 95% confidence intervals per SNR range for individual classifiers and the ensemble.
Figure 3: WuW F1-score (positive class) over the [−10,45] dB SNR range vs. RTF on the Pixel XL.
Figure 4: WuW F1-score (positive class) per SNR range for the best-performing classifiers.
Recent ASR development has placed growing emphasis on generalization across diverse domains and acoustic conditions. Existing approaches typically adapt pretrained ASR models to front-end functions such as wake-up word (WuW) detection through additional training or task-specific modules. In this work, we explore the use of a shared pretrained ASR backbone for WuW detection without gradient-based fine-tuning and examine whether a compact encoder can be extracted using the PCA-based structured pruning approach of SliceGPT. Experiments with Parakeet-TDT-0.6B-v3 and Moonshine-base show that WuW detection performance remains relatively stable when the encoder channel dimension is reduced by 50%. These results suggest that task-relevant compact encoders can be derived from pretrained ASR models without fine-tuning.
In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only triggers for the target user, which helps reduce the computational cost and battery consumption, especially in scenarios where a keyword detector is unpreferable. We achieve this by training a VAD-alike neural network that is conditioned on the target speaker embedding or the speaker verification score. For each frame, personal VAD outputs the probabilities for three classes: non-speech, target speaker speech, and non-target speaker speech. Under our optimal setup, we are able to train a model with only 130K parameters that outperforms a baseline system where individually trained standard VAD and speaker recognition networks are combined to perform the same task.
User-defined keyword spotting (UD-KWS) enables zero-shot wake-word detection from text, but existing systems learn speaker-invariant representations that cannot reject impostors uttering the correct keyword. We address this dual zero-shot setting -- unseen keywords and unseen speakers -- with ZP-KWS, a lightweight framework combining a phoneme-supervised audio encoder with a GE2E-pretrained compact speaker encoder (about 0.9M parameters). Multiplicative late fusion at inference grants each branch independent veto power, supporting modes from conventional detection to strict speaker-gated activation without retraining. On LibriPhrase, Google Speech Commands, and Qualcomm datasets, ZP-KWS reduces target-only FRR at 1% FAR by up to 60% relative to the strongest baseline while maintaining competitive keyword detection, all within a 1.55M parameter budget for edge deployment.
Ming-Hsiang Hu, Kuan-Tang Huang, Chien-Chun Wang +2
Dept. Computer Science and Information Engineering, National Taiwan Normal University, Taiwan · E.SUN Financial Holding Co., Ltd., Taiwan · United Link Co., Ltd., Taiwan