Voice-based interfaces rely on a wake-up word mechanism to initiate communication with devices. However, achieving a robust, energy-efficient, and fast detection remains a challenge. This paper addresses these real production needs by enhancing data with temporal alignments and using detection based on two phases with multi-resolution. It employs two models: a lightweight on-device model for real-time processing of the audio stream and a verification model on the server-side, which is an ensemble of heterogeneous architectures that refine detection. This scheme allows the optimization of two operating points. To protect privacy, audio features are sent to the cloud instead of raw audio. The study investigated different parametric configurations for feature extraction to select one for on-device detection and another for the verification model. Furthermore, thirteen different audio classifiers were compared in terms of performance and inference time. The proposed ensemble outperforms our stronger classifier in every noise condition.
Figures & tables
Figure 1: Attention-based ensemble architecture. Each front-end is the frozen initial stack of a pre-trained classifier and extracts features at its own temporal resolution ri . Projections align the embeddings to a common temporal length tbins′ , which self-attention then fuses. Front-end 1 is the always-on first-stage model, so its states are reused at no additional cost. Generalizes to N front-ends, shown dashed; we use N=2 .
Model
Params.
Oper.
Size (MB)
cnn-fat2019
5.2M
1196.8M
43.6
cnn-trad-pool2
69.5k
10.1M
0.9
gru-att
643.9k
21.9M
5.4
gru-max
145.6k
21.5M
0.8
sgru
145.6k
144.4k
0.8
resnet15
237.4k
456.6M
10.6
Table 1: Parameters, number of operations (multiplications and additions) and size of WuW detection models.
Figure 2: Impact of temporal resolution and number of MFCCs on the WuW F1-score (positive class) by SNR range.
Features
Size
Inference time (ms)
13, w=100, h=50
(29, 13)
25.08
13, w=100, h=20
(71, 13)
51.62
13, w=30, h=10
(148, 13)
73.43
13, w=20, h=10
(149, 13)
82.75
40, w=30, h=10
(148, 40)
83.97
Table 2: MFCC size and inference time on the Pixel XL Android device. The features column shows coefficients, window size ( w ) and hop size ( h ) in milliseconds.
SNR[dB]
device-sgru
resnet15
ensemble
WuW f1-score
95% CI
WuW f1-score
95% CI
WuW f1-score
95% CI
[15, 20]
0.988
0.977, 0.996
0.991
0.981, 0.998
0.989
0.976, 0.996
[10, 15]
0.984
0.969, 0.994
0.993
0.984, 0.999
0.988
0.979, 0.995
[5, 10]
0.978
0.961, 0.990
0.991
0.978, 0.999
0.993
0.989, 0.997
[0, 5]
0.956
0.934, 0.973
0.982
0.967, 0.992
0.983
0.972, 0.989
[-5, 0]
0.888
0.841, 0.923
0.944
0.918, 0.965
0.946
0.929, 0.968
Table 3: WuW F1-score (positive class) and the 95% confidence intervals per SNR range for individual classifiers and the ensemble.
Figure 3: WuW F1-score (positive class) over the [−10,45] dB SNR range vs. RTF on the Pixel XL.
Figure 4: WuW F1-score (positive class) per SNR range for the best-performing classifiers.
Dept. Computer Science and Information Engineering, National Taiwan Normal University, Taiwan · E.SUN Financial Holding Co., Ltd., Taiwan · United Link Co., Ltd., Taiwan