eess.ASSep 24, 2026

Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement

Authors: Clément Laroche, Rasmus Kongsgaard Olsson

Organizations: GN A/S, Audio Research, Lauptrupbjerg 7, 2750, Denmark

Abstract

Deploying real-time speech enhancement on resource-constrained devices requires meeting strict latency, memory, and energy constraints. Microcontroller NPUs can accelerate neural inference under these constraints, but only through a restricted set of operators in static, integer-quantized graphs. Recent speech-enhancement networks have reduced parameter counts and MACs to levels nominally suitable for microcontrollers, but their operators and execution patterns often remain incompatible with restricted NPUs. We address this gap by redesigning LiSenNet, a 37k parameter sub-band dual-path model, for the STM32N6570-DK Neural-ART accelerator. We replace its recurrent bottleneck with convolutional frequency and temporal mixers, reformulate unsupported operations as static int8-compatible primitives, and use bounded decoder activations to preserve quality after quantization. On VoiceBank-DEMAND, the final NPU-compatible model matches or exceeds the recurrent LiSenNet baseline, reaching PESQ 3.08 versus 3.01 in FP32 and 3.01 versus 2.93 in int8. Deployed on a microcontroller, it processes each 16 ms input hop in 4.83 ms, corresponding to a real-time factor of 0.30. Stateless receptive-field recomputation is an order of magnitude slower at the same frame rate despite higher accelerator utilization. These results show that parameter count and operator compatibility, quantization range, and persistent streaming state must be co-designed to achieve efficient real-time speech enhancement on restricted NPUs.

Figures & tables

Explore similar work

CardsList
  1. One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications

    Jun 24, 2026Szu-Wei Fu, Rong Chao, Xuesong Yang +4Speech EnhancementReal-Time Systems

  2. faster-enhancer.c: A Dependency-Free int8 Runtime for Streaming Speech Enhancement on Commodity CPUs

    Jul 28, 2026Gyeongmin KimSpeech EnhancementRuntime

  3. Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement

    Sep 24, 2026Clément Laroche, Riccardo MicciniSpeech EnhancementNeural Processing Units