Deploying real-time speech enhancement on resource-constrained devices requires meeting strict latency, memory, and energy constraints. Microcontroller NPUs can accelerate neural inference under these constraints, but only through a restricted set of operators in static, integer-quantized graphs. Recent speech-enhancement networks have reduced parameter counts and MACs to levels nominally suitable for microcontrollers, but their operators and execution patterns often remain incompatible with restricted NPUs. We address this gap by redesigning LiSenNet, a 37k parameter sub-band dual-path model, for the STM32N6570-DK Neural-ART accelerator. We replace its recurrent bottleneck with convolutional frequency and temporal mixers, reformulate unsupported operations as static int8-compatible primitives, and use bounded decoder activations to preserve quality after quantization. On VoiceBank-DEMAND, the final NPU-compatible model matches or exceeds the recurrent LiSenNet baseline, reaching PESQ 3.08 versus 3.01 in FP32 and 3.01 versus 2.93 in int8. Deployed on a microcontroller, it processes each 16 ms input hop in 4.83 ms, corresponding to a real-time factor of 0.30. Stateless receptive-field recomputation is an order of magnitude slower at the same frame rate despite higher accelerator utilization. These results show that parameter count and operator compatibility, quantization range, and persistent streaming state must be co-designed to achieve efficient real-time speech enhancement on restricted NPUs.
Figures & tables
Model
Params
RF
P-32
P-8
LiSenNet
36.8k
∞
3.01
2.93
+ dual-path conv. mixer
41.1k
68
2.97
2.86
+ NPU-friendly ops, C=20
25.7k
68
2.90
2.85
+ NPU-friendly ops, C=24
36.3k
68
3.01
3.00
+ NPU-friendly ops, C=28
48.7k
68
2.93
2.87
+ dilation 16 (from C=24 )
37.7k
132
3.03
2.95
Table 1: Speech-enhancement ablation on VoiceBank-DEMAND. P-32 and P-8 denote float32 and int8 PESQ, respectively. Each ‘+” row adds the indicated change to the configuration above; C is the bottleneck width. ‘NPU-friendly ops” refers to Sec. 2.2 .
System
STOI
SI-SDR
SIG
BAK
OVRL
Noisy
0.921
8.4
3.04
2.17
1.98
LiSenNet, int8
0.934
17.7
3.02
3.63
2.64
LiSenNet-NPU, int8
0.934
16.9
3.05
3.67
2.69
Table 2: Additional speech-enhancement results on VoiceBank-DEMAND for the int8 models in Table 1 . SI-SDR is reported in dB. SIG, BAK, and OVRL are DNSMOS scores.
Model
Params
RF
MACs/frame
State tensor
State size
Weights
ms/frame
RTF
P-8
LiSenNet + temporal conv. mixer + frequency conv. mixer
+ NPU-friendly ops, C=20
25.7 k
68
0.92 M
17
47 KiB
24.9 KiB
2.59
0.16
2.85
NPU-friendly ops, C=24
36.3 k
68
1.30 M
17
57 KiB
35.3 KiB
2.79
0.17
2.96
NPU-friendly ops, C=28
48.7 k
68
1.75 M
17
66 KiB
47.3 KiB
3.15
0.20
2.87
+ dilation 16 (from C=24 )
37.7 k
132
1.38 M
19
105 KiB
36.6 KiB
3.63
0.23
2.95
+ third DPC block
46.2 k
196
1.69 M
25
154 KiB
44.8 KiB
4.88
0.30
2.99
Table 3: Compute footprint, measured latency, and deployed int8 PESQ (P-8) on the STM32N6 for LiSenNet-NPU, its ablations, and literature baselines. As in Table 1 , each “+” row retains all modifications above it and adds the indicated change.
Different real-time speech applications impose distinct latency budgets, often requiring separately trained enhancement models for each scenario. In this paper, we propose a one-for-all, real-time universal speech enhancement model that provides explicit control over both algorithmic and computational latency. Algorithmic latency is flexibly adjusted via configurable look-ahead frames. To avoid learning inefficiency caused by varying padding configurations, we introduce parallel convolutional layers corresponding to different look-ahead settings. Computational latency is controlled through an early-exit mechanism, enabling inference at different network depths. To narrow the performance gap between specialized and flexible models, we propose a two-stage training strategy with a shared-to-multiple decoder transition. Overall, the proposed framework enables a single model to be deployed across diverse latency budgets without retraining separate models. Model weights are available for download at: https://huggingface.co/nvidia/Real-time_RE-USE
This is an implementation and measurement study of what it costs to run a streaming speech enhancer on a CPU. We port FastEnhancer-Medium at 48 kHz to faster-enhancer.c, a C runtime with six int8 GEMM tiers selected at initialization, leaving architecture and weights untouched. One Apple M2 core reaches 0.069 real-time factor, against 0.230 for the fp32 ONNX Runtime graph on the same machine, a 3.3x speedup. A Galaxy S23+ (Snapdragon 8 Gen 2) reaches 0.096. The speedup comes from specializing every layer of the runtime around one fixed model. Activation ranges are recomputed per frame, so no calibration set is needed; the k=3 convolutions use Winograd F(2,3); cross-stage state is fp16; the GRU and the dequantization epilogues are fused; and nothing is allocated after startup. Over 824 VoiceBank-DEMAND utterances the engine tracks fp32 to within -0.006 PESQ and -0.08 dB SNR. Speed alone does not settle deployment cost. The enhancer holds a fraction of a core for as long as the microphone is open, so its real-time factor is a duty cycle. A benchmark races through a file; an audio callback does not. Pacing to the 6.67 ms deadline costs 4.2x per frame, saves 49% of the energy, and leaves the cheapest core placement missing 96% of its deadlines. All SIMD tiers within an architecture family emit byte-identical output. The runtime is released as a dependency-free library.
Gyeongmin Kim
Department of Computer Science, Hanyang University, Seoul, Korea
Deep learning-based speech enhancement is increasingly deployed on-device in hearing aids, headsets, and earbuds. Most of these devices, however, can only accelerate static int8 graphs, so a depth-varying network must be implemented as several graphs, orchestrated by a policy. In this paper, we supervise every intermediate depth of one causal model, then we fine-tune its output heads to guarantee that deeper outputs are never worse than shallower ones. Using this training protocol, we can derive a family of static models that are more Pareto-efficient than their equivalently-sized counterparts trained from scratch on the same budget. Specifically, we achieve up to 0.11 higher PESQ for equivalent compute, and match the best PESQ at 30% less compute. We then quantize the models to int8 and measure the latency-quality frontier on an STM32N6 microcontroller. On VoiceBank-DEMAND, the dynamic enhancer lies on the same frontier as the static models, rather than trading quality for dynamic execution. Running the policy on the companion Cortex-M55 takes only 26 μs per frame, while splitting the enhancer into separate NPU graphs adds 2.2% latency overhead. The cost of dynamic execution is therefore small.
Clément Laroche, Riccardo Miccini
GN A/S, Denmark · Technical University of Denmark, Denmark