Accurate, efficient, and low-latency spatial beamforming is a key component in emerging smart hearable devices, enhancing speech while suppressing noise and interference. However, handling multiple input sources under strict real-time constraints poses significant challenges for the low-power, resource-constrained microcontroller units (MCUs) used in hearables. We present an optimized methodology for the real-time execution of a neural-network-based minimum variance distortionless response (MVDR) beamformer on MCUs. Using six microphones and a three-stage mixed-precision scheme (float32 MVDR, int8 CNN, float16 Transformer), the pipeline pairs a CNN that estimates the MVDR weights with a lightweight Transformer that applies a per-frame correction. By time-slicing weight estimation with beamforming, it achieves a 15ms per-frame latency while refreshing a complete set of CNN-derived weights every 564ms. The deployed mixed-precision pipeline attains a short-time objective intelligibility (STOI) of 97.65%, a scale-invariant signal-to-noise ratio (SI-SNR) of 20.26dB, and a wideband PESQ of 3.676 at an average power of 45.9mW. A speech activity detection (SAD) module (98.5% accuracy, 0.62mJ per inference) bypasses the pipeline during silence; under realistic deployment conditions, the system exceeds the 16h all-day target on a 100mAh battery, with an estimated lifetime of up to ∼20h. To our knowledge, this is the first real-time multi-channel Transformer-based neural beamforming pipeline deployed on an MCU-class device.
Fig. 3 : Simplified architecture of the GAP9 SoC in the investigated configuration.
Fig. 4 : Experimental setup
Model
STOI (%)
ESTOI (%)
PESQ-WB
SI-SNR (dB)
Noisy
92.38 ± 0.28
86.97 ± 0.33
2.609 ± 0.034
10.15 ± 0.13
CNN
94.19 ± 0.23
89.95 ± 0.47
2.947 ± 0.041
17.70 ± 0.31
Transformer
95.98 ± 0.12
91.96 ± 0.17
3.262 ± 0.013
19.79 ± 0.10
CNN + Transformer
98.04 ± 0.07
95.52 ± 0.12
3.747 ± 0.018
22.02 ± 0.12
Table I : Speech enhancement performance of the CNN, Transformer, and hybrid Transformer + CNN architectures, with the unprocessed noisy input as reference. Values are mean ± std.
Implementation
Cycles
Attention buffers
NNTool baseline, full attention matrix
9.25M
942 kB L2 + 400 kB L3
Sequential heads, one full matrix per head
3.70M
132 kB L2
Row-wise attention in L1
2.22M
L1 only
Fused row-wise attention kernel
1.65M
L1 only
Table II : Transformer attention optimizations.
Base
Optimized
(G1 @ 0.8 V, 370 MHz) (G2 @ 0.8 V, 370 MH)
(G1 @ 0.65 V, 240 MHz) (G2 @ 0.8 V, 370 MH)
G1 execution time [ms]
244
379
STFT ( fp32 , 2.4 M cyc.)
6.5
10
Mask est. ( int8 , NE16)
203
315
MVDR ( fp32 , 13 M cyc.)
35
54
G2 execution time ( fp16,fp32 , 1.85 M cyc.)
5
5
Table III : Timing and energy comsumption on GAP9 under interleaved execution. We report for G2 the total energy consumed by all executions in the time it takes for a G1’s inference to be completed.
Fig. 7 : Power profile of our application.
Model
Mpar
Quantization
QAT
Device
Deployment
ms/inf
MAC/cyc
GMAC/J
STOI
TinyLSTM [ 12 ]
0.33
INT8
yes
STM32F746VE
N/A
4.26
0.36
0.14
N/A
TinyLSTM [ 12 ]
0.46
INT8
yes
STM32F746VE
N/A
2.39
0.36
0.14
N/A
RNNoise [ 14 ]
0.21
INT8
yes
STM32L476
NNoM w/ CMSIS-NN
3.28
0.45
1.84
N/A
LSTM256 [ 15 ]
1.24
MixFP16-INT8
no
9-core RISC-V
GAPFlow
2.50
2.11
17.78
N/A
GRU256 [ 15 ]
0.98
MixFP16-INT8
no
9-core RISC-V
GAPFlow
1.70
2.41
17.46
N/A
TCN ts=4 [ 37 ]
0.83
INT8-BFP16
no
9-core RISC-V
GAPFlow
4.3
3.3
28.90
93.36
Table IV : Comparison with other MCU-deployed SE pipelines.
Real-time binaural speech enhancement is constrained by latency, computational cost, and inter-device communication, yet existing efficient solutions predominantly address single-channel settings. In this paper, we introduce RT-Tango, a real-time distributed binaural speech enhancement framework designed for streaming on resource-constrained platforms and specifically for hearing aids. RT-Tango relies on a two-stage distributed architecture combining perceptually motivated ERB feature compression, lightweight grouped recurrent mask estimation, and temporal sparsification to reduce computational cost. Stringent latency constraints are addressed by decoupling spectral resolution from algorithmic delay using an asymmetric STFT, together with causal recurrent inference and online estimation of spatial statistics. Experimental results show that RT-Tango achieves competitive speech enhancement while significantly reducing MACs operations and functioning at ultra-low latencies as low as 8 ms.
Z. Benslimane, P. Chouteau, M. Poreba +4
Universit´e Paris-Saclay, CEA, List, F-91120 Palaiseau, France · Universit´e de Lorraine, CNRS, Inria, LORIA, F-54000 Nancy, France
Deploying real-time speech enhancement on resource-constrained devices requires meeting strict latency, memory, and energy constraints. Microcontroller NPUs can accelerate neural inference under these constraints, but only through a restricted set of operators in static, integer-quantized graphs. Recent speech-enhancement networks have reduced parameter counts and MACs to levels nominally suitable for microcontrollers, but their operators and execution patterns often remain incompatible with restricted NPUs. We address this gap by redesigning LiSenNet, a 37k parameter sub-band dual-path model, for the STM32N6570-DK Neural-ART accelerator. We replace its recurrent bottleneck with convolutional frequency and temporal mixers, reformulate unsupported operations as static int8-compatible primitives, and use bounded decoder activations to preserve quality after quantization. On VoiceBank-DEMAND, the final NPU-compatible model matches or exceeds the recurrent LiSenNet baseline, reaching PESQ 3.08 versus 3.01 in FP32 and 3.01 versus 2.93 in int8. Deployed on a microcontroller, it processes each 16 ms input hop in 4.83 ms, corresponding to a real-time factor of 0.30. Stateless receptive-field recomputation is an order of magnitude slower at the same frame rate despite higher accelerator utilization. These results show that parameter count and operator compatibility, quantization range, and persistent streaming state must be co-designed to achieve efficient real-time speech enhancement on restricted NPUs.
Clément Laroche, Rasmus Kongsgaard Olsson
GN A/S, Audio Research, Lauptrupbjerg 7, 2750, Denmark
The minimum variance distortionless response (MVDR) beamformer is widely used for multichannel speech enhancement due to strong noise suppression while preserving target signals. In practice, its performance is sensitive to microphone self-noise and array mismatches. Existing approaches typically rely on fixed, manually tuned WNG thresholds or diagonal loading, leading to suboptimal performance under unknown or time-varying acoustic conditions. This paper proposes a data-driven MVDR framework that adaptively estimates the WNG constraint using a deep neural network. The network jointly predicts a time-frequency noise mask for covariance estimation and a frequency-dependent WNG threshold, enabling dynamic robustness-directivity control. A differentiable robust MVDR layer is integrated into the framework, allowing end-to-end optimization. Experiments demonstrate consistent improvements in speech quality and intelligibility over conventional fixed-WNG MVDR methods.
Yongyi Deng, Hanchen Pei, Jianbo Ma +3
School of Electronic Information, Wuhan University, Wuhan, Hubei, China · Dolby Laboratories · CIAIC, Northwestern Polytechnical University, Xi’an, Shaanxi, China +1