Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement
Organizations: GN A/S, Denmark · Technical University of Denmark, Denmark
Abstract
Deep learning-based speech enhancement is increasingly deployed on-device in hearing aids, headsets, and earbuds. Most of these devices, however, can only accelerate static int8 graphs, so a depth-varying network must be implemented as several graphs, orchestrated by a policy. In this paper, we supervise every intermediate depth of one causal model, then we fine-tune its output heads to guarantee that deeper outputs are never worse than shallower ones. Using this training protocol, we can derive a family of static models that are more Pareto-efficient than their equivalently-sized counterparts trained from scratch on the same budget. Specifically, we achieve up to 0.11 higher PESQ for equivalent compute, and match the best PESQ at 30% less compute. We then quantize the models to int8 and measure the latency-quality frontier on an STM32N6 microcontroller. On VoiceBank-DEMAND, the dynamic enhancer lies on the same frontier as the static models, rather than trading quality for dynamic execution. Running the policy on the companion Cortex-M55 takes only 26 s per frame, while splitting the enhancer into separate NPU graphs adds 2.2% latency overhead. The cost of dynamic execution is therefore small.
Figures & tables
| Topology | MMAC/s | ms/frame | s/MMAC |
|---|---|---|---|
| , | 5.73 | 0.640 | 112 |
| , | 7.30 | 0.751 | 103 |
| , | 7.30 | 1.001 | 137 |
| , | 8.86 | 0.852 | 96 |
| , | 10.43 | 1.215 | 117 |
| , | 13.57 | 1.427 | 105 |
| Model | Variant | MMAC/s | int8 PESQ | ms/frame |
|---|---|---|---|---|
| Within a fixed ladder | ||||
| Static | fixed | 10.4 | 2.716 | 1.215 |
| Routed | 8.58 | 2.694 | 0.930 | |
| Static | fixed | 29.3 | 2.781 | 2.472 |
| Routed | 20.19 | 2.773 | 1.668 | |
| Across static and dynamic topologies | ||||