Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must meet high accuracy requirements while operating under strict constraints on computational power, memory footprint, and real-time latency. In this work, we present an application of the STMC (Short-Term Memory Convolutions) framework to adapt a modular CNN model for online, LSTM-like inference. Our approach reduces power consumption and redundant computations while maintaining the stability and simplicity of training CNNs. We achieve up to 82% and 46% MCPS reduction compared to equivalently frequent standard CNN execution and vanilla STMC, respectively. The best configuration achieves 93.8% accuracy on the 11-class Google Speech Commands task and 97.1% on the same task with zero-padded data.
Figures & tables
Figure 1: New states after pooling layer with the stride of 2.
Figure 2: Difference between STMC (left) and SM-STMC (right).
Model
Emb
Cls
Total
SM-STMC 1
513 + 1776 + 6496
18443
27228
SM-STMC 2
105 + 816 + 1472 + 12064
18443
32900
LSTM
20992
6155
27147
Table 1: Parameters of the backbone’s embedding and classifier.
Model
Total buffer size
STMC 1
17184
SM-STMC 1
9280
STMC 2
21632
SM-STMC 2
7520
Table 2: Buffer size (int8).
Model
Precision
Recall
F1-score
VGG 1 ( 1×1 s )
0.9362
0.9278
0.9292
SM-STMC 1
0.9494
0.9382
0.9403
VGG 2 ( 1×1 s )
0.9159
0.9035
0.9055
SM-STMC 2
0.9322
0.9162
0.9188
LSTM
0.9569
0.9511
0.9521
Table 3: Weighted average performance comparison on the standard 1 s GSC dataset.
Model
Precision
Recall
F1-score
VGG 1 ( 1×1 s )
0.8353
0.4981
0.5798
VGG 1 ( 8×1 s )
0.9733
0.9708
0.9712
SM-STMC 1
0.9738
0.9710
0.9715
VGG 2 ( 1×1 s )
0.7790
0.6182
0.6386
VGG 2 ( 8×1 s )
0.9586
0.9524
0.9532
SM-STMC 2
0.9611
0.9552
0.9561
Table 4: Weighted average performance comparison on the extended 2 s GSC dataset with silence added.
Figure 3: Influence of temporal shift on accuracy. VGG and STMC accuracy on a 2-second zero-padded “stop” sample.
Model
Emb
Cls
Memory
Total
VGG 1 ( 1×1 s )
7.40
0.02
–
7.42
VGG 1 ( 8×1 s )
59.20
0.16
–
59.36
STMC 1
16.12
0.16
0.56
16.84
SM-STMC 1
10.82
0.16
0.39
11.37
VGG 2 ( 1×1 s )
3.66
0.02
–
3.68
VGG 2 ( 8×1 s )
29.28
0.16
–
29.44
Table 5: Computational cost comparison of embedding, classifier, and memory management in MCPS.
Model
Comp. Latency [ms]
Avg Latency [ms]
VGG 1
101
601 ± 500
STMC 1
4
-156 ± 8
SM-STMC 1
10
-146 ± 32
VGG 2
50
550 ± 500
STMC 2
2
-158 ± 8
SM-STMC 2
9
-116 ± 64
Table 6: Computational latency and average latency of the models measured from the end of the uttered keyword to its first corresponding prediction.
Dept. Computer Science and Information Engineering, National Taiwan Normal University, Taiwan · E.SUN Financial Holding Co., Ltd., Taiwan · United Link Co., Ltd., Taiwan