Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must meet high accuracy requirements while operating under strict constraints on computational power, memory footprint, and real-time latency. In this work, we present an application of the STMC (Short-Term Memory Convolutions) framework to adapt a modular CNN model for online, LSTM-like inference. Our approach reduces power consumption and redundant computations while maintaining the stability and simplicity of training CNNs. We achieve up to 82% and 46% MCPS reduction compared to equivalently frequent standard CNN execution and vanilla STMC, respectively. The best configuration achieves 93.8% accuracy on the 11-class Google Speech Commands task and 97.1% on the same task with zero-padded data.
Figures & tables
Figure 1: New states after pooling layer with the stride of 2.
Figure 2: Difference between STMC (left) and SM-STMC (right).
Model
Emb
Cls
Total
SM-STMC 1
513 + 1776 + 6496
18443
27228
SM-STMC 2
105 + 816 + 1472 + 12064
18443
32900
LSTM
20992
6155
27147
Table 1: Parameters of the backbone’s embedding and classifier.
Model
Total buffer size
STMC 1
17184
SM-STMC 1
9280
STMC 2
21632
SM-STMC 2
7520
Table 2: Buffer size (int8).
Model
Precision
Recall
F1-score
VGG 1 ( 1×1 s )
0.9362
0.9278
0.9292
SM-STMC 1
0.9494
0.9382
0.9403
VGG 2 ( 1×1 s )
0.9159
0.9035
0.9055
SM-STMC 2
0.9322
0.9162
0.9188
LSTM
0.9569
0.9511
0.9521
Table 3: Weighted average performance comparison on the standard 1 s GSC dataset.
Model
Precision
Recall
F1-score
VGG 1 ( 1×1 s )
0.8353
0.4981
0.5798
VGG 1 ( 8×1 s )
0.9733
0.9708
0.9712
SM-STMC 1
0.9738
0.9710
0.9715
VGG 2 ( 1×1 s )
0.7790
0.6182
0.6386
VGG 2 ( 8×1 s )
0.9586
0.9524
0.9532
SM-STMC 2
0.9611
0.9552
0.9561
Table 4: Weighted average performance comparison on the extended 2 s GSC dataset with silence added.
Figure 3: Influence of temporal shift on accuracy. VGG and STMC accuracy on a 2-second zero-padded “stop” sample.
Model
Emb
Cls
Memory
Total
VGG 1 ( 1×1 s )
7.40
0.02
–
7.42
VGG 1 ( 8×1 s )
59.20
0.16
–
59.36
STMC 1
16.12
0.16
0.56
16.84
SM-STMC 1
10.82
0.16
0.39
11.37
VGG 2 ( 1×1 s )
3.66
0.02
–
3.68
VGG 2 ( 8×1 s )
29.28
0.16
–
29.44
Table 5: Computational cost comparison of embedding, classifier, and memory management in MCPS.
Model
Comp. Latency [ms]
Avg Latency [ms]
VGG 1
101
601 ± 500
STMC 1
4
-156 ± 8
SM-STMC 1
10
-146 ± 32
VGG 2
50
550 ± 500
STMC 2
2
-158 ± 8
SM-STMC 2
9
-116 ± 64
Table 6: Computational latency and average latency of the models measured from the end of the uttered keyword to its first corresponding prediction.
Keyword spotting (KWS) models on embedded devices often need to add new keywords after deployment, but updates are difficult when original training data are unavailable and regressions on existing triggers are unacceptable. At a fixed operating point, our method reduces average new-keyword false reject rate (FRR) from 6.46 to 4.37 versus a parameter-matched separate-model baseline and outperforms parameter-efficient tuning baselines (adapters, LoRA), while using fewer multiply-accumulate operations (MACs) under the same added-parameter budget (≤10k): 16.34M vs 18.45M/20.52M. We achieve this via parameter-capped modular expansion: the base network, including batch-normalization statistics and the core classifier, is frozen, and only a lightweight expansion branch with a separate new-keyword head is trained, preserving core logits, shipped outputs, and thresholds for existing keywords.
Viktor Khaymonenko, Dzmitry Saladukha, Aliaksei Rak +1
Keyword spotting (KWS), the task of identifying predefined words in speech, is a core capability of voice-enabled devices. Achieving high KWS accuracy under tight parameter budgets across different vocabulary sizes remains challenging. We present CircleMatch, a matching framework enabling KWS with very few parameters. Its encoder independently compresses frequency bands and fuses them into frame features. These features are then matched against learned class-specific prototypes to produce temporal response curves. Parameter-free circular aggregation encodes time as angles and summarizes response distributions and relative timing for classification. We develop four tiny variants, Circle-D4, Circle-D8, Circle-D16, and Circle-D32, ranging from approximately 1k to 7k parameters in the 12-class setting. Experiments with multiple random seeds on Speech Commands v1/v2 and the English and Spanish Micro subsets of the Multilingual Spoken Words Corpus demonstrate competitive accuracy with tiny models. Our qualitative analysis further suggests approximate shift equivariance of prototype responses and adaptation to temporal compression. Code and model weights are available at https://github.com/ora942878/CircleMatch.
User-defined keyword spotting (UD-KWS) enables zero-shot wake-word detection from text, but existing systems learn speaker-invariant representations that cannot reject impostors uttering the correct keyword. We address this dual zero-shot setting -- unseen keywords and unseen speakers -- with ZP-KWS, a lightweight framework combining a phoneme-supervised audio encoder with a GE2E-pretrained compact speaker encoder (about 0.9M parameters). Multiplicative late fusion at inference grants each branch independent veto power, supporting modes from conventional detection to strict speaker-gated activation without retraining. On LibriPhrase, Google Speech Commands, and Qualcomm datasets, ZP-KWS reduces target-only FRR at 1% FAR by up to 60% relative to the strongest baseline while maintaining competitive keyword detection, all within a 1.55M parameter budget for edge deployment.
Ming-Hsiang Hu, Kuan-Tang Huang, Chien-Chun Wang +2
Dept. Computer Science and Information Engineering, National Taiwan Normal University, Taiwan · E.SUN Financial Holding Co., Ltd., Taiwan · United Link Co., Ltd., Taiwan