Deploying real-time speech enhancement on resource-constrained devices requires meeting strict latency, memory, and energy constraints. Microcontroller NPUs can accelerate neural inference under these constraints, but only through a restricted set of operators in static, integer-quantized graphs. Recent speech-enhancement networks have reduced parameter counts and MACs to levels nominally suitable for microcontrollers, but their operators and execution patterns often remain incompatible with restricted NPUs. We address this gap by redesigning LiSenNet, a 37k parameter sub-band dual-path model, for the STM32N6570-DK Neural-ART accelerator. We replace its recurrent bottleneck with convolutional frequency and temporal mixers, reformulate unsupported operations as static int8-compatible primitives, and use bounded decoder activations to preserve quality after quantization. On VoiceBank-DEMAND, the final NPU-compatible model matches or exceeds the recurrent LiSenNet baseline, reaching PESQ 3.08 versus 3.01 in FP32 and 3.01 versus 2.93 in int8. Deployed on a microcontroller, it processes each 16 ms input hop in 4.83 ms, corresponding to a real-time factor of 0.30. Stateless receptive-field recomputation is an order of magnitude slower at the same frame rate despite higher accelerator utilization. These results show that parameter count and operator compatibility, quantization range, and persistent streaming state must be co-designed to achieve efficient real-time speech enhancement on restricted NPUs.
Figures & tables
Model
Params
RF
P-32
P-8
LiSenNet
36.8k
∞
3.01
2.93
+ dual-path conv. mixer
41.1k
68
2.97
2.86
+ NPU-friendly ops, C=20
25.7k
68
2.90
2.85
+ NPU-friendly ops, C=24
36.3k
68
3.01
3.00
+ NPU-friendly ops, C=28
48.7k
68
2.93
2.87
+ dilation 16 (from C=24 )
37.7k
132
3.03
2.95
Table 1: Speech-enhancement ablation on VoiceBank-DEMAND. P-32 and P-8 denote float32 and int8 PESQ, respectively. Each ‘+” row adds the indicated change to the configuration above; C is the bottleneck width. ‘NPU-friendly ops” refers to Sec. 2.2 .
System
STOI
SI-SDR
SIG
BAK
OVRL
Noisy
0.921
8.4
3.04
2.17
1.98
LiSenNet, int8
0.934
17.7
3.02
3.63
2.64
LiSenNet-NPU, int8
0.934
16.9
3.05
3.67
2.69
Table 2: Additional speech-enhancement results on VoiceBank-DEMAND for the int8 models in Table 1 . SI-SDR is reported in dB. SIG, BAK, and OVRL are DNSMOS scores.
Model
Params
RF
MACs/frame
State tensor
State size
Weights
ms/frame
RTF
P-8
LiSenNet + temporal conv. mixer + frequency conv. mixer
+ NPU-friendly ops, C=20
25.7 k
68
0.92 M
17
47 KiB
24.9 KiB
2.59
0.16
2.85
NPU-friendly ops, C=24
36.3 k
68
1.30 M
17
57 KiB
35.3 KiB
2.79
0.17
2.96
NPU-friendly ops, C=28
48.7 k
68
1.75 M
17
66 KiB
47.3 KiB
3.15
0.20
2.87
+ dilation 16 (from C=24 )
37.7 k
132
1.38 M
19
105 KiB
36.6 KiB
3.63
0.23
2.95
+ third DPC block
46.2 k
196
1.69 M
25
154 KiB
44.8 KiB
4.88
0.30
2.99
Table 3: Compute footprint, measured latency, and deployed int8 PESQ (P-8) on the STM32N6 for LiSenNet-NPU, its ablations, and literature baselines. As in Table 1 , each “+” row retains all modifications above it and adds the indicated change.