Transformer-based Neural Beamforming for Real-Time Speech Enhancement on Smart Low-Power Hearable Devices
Organizations: Electrical, Electronic and Information Engineering (DEI), University of Bologna, Italy
Abstract
Accurate, efficient, and low-latency spatial beamforming is a key component in emerging smart hearable devices, enhancing speech while suppressing noise and interference. However, handling multiple input sources under strict real-time constraints poses significant challenges for the low-power, resource-constrained microcontroller units (MCUs) used in hearables. We present an optimized methodology for the real-time execution of a neural-network-based minimum variance distortionless response (MVDR) beamformer on MCUs. Using six microphones and a three-stage mixed-precision scheme (float32 MVDR, int8 CNN, float16 Transformer), the pipeline pairs a CNN that estimates the MVDR weights with a lightweight Transformer that applies a per-frame correction. By time-slicing weight estimation with beamforming, it achieves a 15ms per-frame latency while refreshing a complete set of CNN-derived weights every 564ms. The deployed mixed-precision pipeline attains a short-time objective intelligibility (STOI) of 97.65%, a scale-invariant signal-to-noise ratio (SI-SNR) of 20.26dB, and a wideband PESQ of 3.676 at an average power of 45.9mW. A speech activity detection (SAD) module (98.5% accuracy, 0.62mJ per inference) bypasses the pipeline during silence; under realistic deployment conditions, the system exceeds the 16h all-day target on a 100mAh battery, with an estimated lifetime of up to 20h. To our knowledge, this is the first real-time multi-channel Transformer-based neural beamforming pipeline deployed on an MCU-class device.
Figures & tables
| Model | STOI (%) | ESTOI (%) | PESQ-WB | SI-SNR (dB) |
| Noisy | 92.38 0.28 | 86.97 0.33 | 2.609 0.034 | 10.15 0.13 |
| CNN | 94.19 0.23 | 89.95 0.47 | 2.947 0.041 | 17.70 0.31 |
| Transformer | 95.98 0.12 | 91.96 0.17 | 3.262 0.013 | 19.79 0.10 |
| CNN + Transformer | 98.04 0.07 | 95.52 0.12 | 3.747 0.018 | 22.02 0.12 |
| Implementation | Cycles | Attention buffers |
| NNTool baseline, full attention matrix | 942 kB L2 + 400 kB L3 | |
| Sequential heads, one full matrix per head | 132 kB L2 | |
| Row-wise attention in L1 | L1 only | |
| Fused row-wise attention kernel | L1 only |
| Base | Optimized | |
| (G1 @ 0.8 V, 370 MHz) (G2 @ 0.8 V, 370 MH) | (G1 @ 0.65 V, 240 MHz) (G2 @ 0.8 V, 370 MH) | |
| G1 execution time [ms] | 244 | 379 |
| STFT ( fp32 , 2.4 M cyc.) | 6.5 | 10 |
| Mask est. ( int8 , NE16) | 203 | 315 |
| MVDR ( fp32 , 13 M cyc.) | 35 | 54 |
| G2 execution time ( fp16,fp32 , 1.85 M cyc.) | 5 | 5 |
| Model | Mpar | Quantization | QAT | Device | Deployment | ms/inf | MAC/cyc | GMAC/J | STOI |
| TinyLSTM [ 12 ] | 0.33 | INT8 | yes | STM32F746VE | N/A | 4.26 | 0.36 | 0.14 | N/A |
| TinyLSTM [ 12 ] | 0.46 | INT8 | yes | STM32F746VE | N/A | 2.39 | 0.36 | 0.14 | N/A |
| RNNoise [ 14 ] | 0.21 | INT8 | yes | STM32L476 | NNoM w/ CMSIS-NN | 3.28 | 0.45 | 1.84 | N/A |
| LSTM256 [ 15 ] | 1.24 | MixFP16-INT8 | no | 9-core RISC-V | GAPFlow | 2.50 | 2.11 | 17.78 | N/A |
| GRU256 [ 15 ] | 0.98 | MixFP16-INT8 | no | 9-core RISC-V | GAPFlow | 1.70 | 2.41 | 17.46 | N/A |
| TCN ts=4 [ 37 ] | 0.83 | INT8-BFP16 | no | 9-core RISC-V | GAPFlow | 4.3 | 3.3 | 28.90 | 93.36 |