Voice-controlled interaction in industrial settings is hampered by acoustic noise, which severely degrades audio-only speech recognition. Audio-visual speech recognition (AVSR) addresses this by fusing lip-motion cues with the audio stream, but state-of-the-art pipelines rely on three-dimensional convolutions, recurrent units, and attention modules that exceed the budget of typical edge devices. We present NAVIR, an end-to-end AVSR system targeting the BrainChip Akida neuromorphic processor, which natively supports only sequential two-dimensional convolutional inference. The pipeline factorises spatial and temporal encoding into separate AkidaNet-based modules: a per-frame visual encoder, a temporal video encoder, and a spectrogram audio encoder, fused by a lightweight predictor head and decoded by constrained beam search. Models are trained with connectionist temporal classification on noise-augmented audio and then fine-tuned with quantization-aware training. On the GRID benchmark, the quantized audio-visual model reaches 14.0% word error rate (WER) under noise on the unseen-speaker split and 3.3% WER on the overlapped-speaker split, against 22.5% and 11.8% for audio-only baselines, and it attains 98.6% command accuracy at 1.5% WER on a task-specific industrial-command corpus. Operation-count analysis indicates a 13-fold energy advantage of the spiking formulation over its artificial neural network counterpart at 27.6% mean firing rate. On-board measurements show roughly 5-fold lower energy per inference than a Raspberry Pi central processing unit on the lip-reading model, and over 100-fold lower than a laptop graphics processing unit, while sustaining 14.5 inferences per second. To the best of our knowledge, this is the first complete multimodal AVSR pipeline running on neuromorphic hardware of this class.
Figures & tables
Model
Year
Setting
WER (%)
LipNet [ 28 ]
2016
Overlapped
4.80
WLAS [ 29 ]
2017
Overlapped
3.00
LCANet [ 30 ]
2018
Overlapped
2.90
LipSound [ 31 ]
2019
Overlapped
2.50
DualLip [ 32 ]
2020
Overlapped
2.71
HLR-Net [ 33 ]
2021
Overlapped
3.30
TABLE I: State-of-the-art lip-reading WER on the GRID corpus. All baselines use full 3D convolutions, recurrent layers or attention, none of which are supported on AKD1000.
Fig. 1: NAVIR audio-visual pipeline. Aligned video and audio windows are independently encoded by AkidaNet-based per-frame, temporal-video and spectrogram-audio encoders. Embeddings are concatenated and fed to an MLP predictor head, whose CTC output is decoded by constrained beam search. All blocks are AKD1000-compatible (no 3D convolutions, recurrence or attention).
GRID
NAVIR
Batch size
16
3
Epochs, float (unseen/overlap)
30 / 100
— / 200
Epochs, QAT (unseen/overlap)
15 / 100
— / 200
Video window (frames)
15
18
Video overlap (frames)
10
12
Audio (spectrogram) window
60
60
TABLE II: Per-dataset training and pre-processing hyperparameters.
Training modalities
Clean audio
Noisy audio
Clean audio
3.7
79.9
Noisy audio
4.4
21.0
Video
34.0
34.0
Noisy audio + video
7.7
16.6
Clean audio + video
3.6
78.6
TABLE III: Non-quantized WER (%) on GRID (unseen-speaker split) under clean and noisy audio test conditions.
Training modalities
Clean audio
Noisy audio
Clean audio
4.2
77.3
Noisy audio
5.2
22.5
Video
35.3
35.3
Noisy audio + video
5.3
14.0
Clean audio + video
3.2
77.8
TABLE IV: Quantized WER (%) on GRID (unseen-speaker split) after QAT.
Training modalities
Clean audio
Noisy audio
Clean audio
1.3
79.8
Noisy audio
1.9
10.4
Video
9.1
9.1
Noisy audio + video
0.7
2.8
Clean audio + video
0.7
77.4
TABLE V: Non-quantized WER (%) on GRID (overlapped-speaker split) under clean and noisy audio test conditions. Both float and QAT models were trained for 100 epochs in this regime.
Training modalities
Clean audio
Noisy audio
Clean audio
2.0
78.7
Noisy audio
2.3
11.8
Video
6.7
6.7
Noisy audio + video
0.8
3.3
Clean audio + video
0.8
77.1
TABLE VI: Quantized WER (%) on GRID (overlapped-speaker split) after QAT.
TABLE VIII: Quantized NAVIR results after QAT: WER (%) / command accuracy (%).
Model
WER (%)
Params
FLOPs
ANN energy
LipNet [ 28 ]
11.4
4.57 M
10.69 G
19.78 mJ
Wu et al. [ 35 ]
10.21
10.89 M
84.62 G
156.54 mJ
Ours (video-only)
35.30
1.49 M
2.24 G
4.15 mJ
TABLE IX: Model complexity and theoretical ANN inference energy on GRID.
Model
ANN energy
SNN energy
ANN/SNN gain
LipNet [ 28 ]
19.78 mJ
—
—
Wu et al. [ 35 ]
156.54 mJ
—
—
Ours (video-only)
4.15 mJ
314.92 μ J
13.17 ×
TABLE X: Theoretical SNN energy and efficiency gain.
Fig. 2: Pareto view of GRID lip-reading models, computational cost (GFLOPs, log scale) versus word error rate, for both overlapped-speaker (red diamonds) and unseen-speaker (blue circles) test conditions. Bubble area is proportional to parameter count. Our quantized video-only model (green) sits alone on the low-FLOPs frontier.
Backend
it/s
mWh@5 min
Δ mWh
mWh/inf.
Pi + AKD1000 (Akida)
14.55
0 462.0
0 72.0
0.0165
Pi + CPU (Akida)
15.55
0 768.0
378.0
0.0810
Pi + CPU (Keras)
1.10
0 531.9
141.9
0.4306
Laptop + GPU (Keras)
2.54
2,569.6
1,289.6
1.6913
TABLE XI: Video-only model: per-inference power consumption across backends. Idle baselines: Pi 390 mWh / 5 min, GPU 1,280 mWh / 5 min. Pi figures cover total system draw (FNB58). GPU figures cover GPU-only draw via nvidia-smi , normalised to a 5-minute window.
Backend
it/s
mWh@5 min
Δ mWh
mWh/inf.
Pi + AKD1000 (Akida)
2.61
0 459.9
0 69.9
0.0894
Pi + CPU (Akida)
13.36
0 737.2
347.2
0.0866
Pi + CPU (Keras)
0.63
0 528.5
138.5
0.7381
Laptop + GPU (Keras)
1.29
2,566.2
1,286.2
3.3354
TABLE XII: Audio-video model: per-inference power consumption across backends. Same idle baselines and measurement scopes as Table XI .
Fig. 3: NAVIR demonstration system. An Obsbot Meet SE webcam supplies frame and audio capture. A Raspberry Pi 5 with an M.2 HAT hosts the BrainChip AKD1000 co-processor. A uFactory xArm 6 robotic arm executes the recognised commands.
Leonidas Delimpasis received the five-year (M.Eng. equivalent) degree in Electrical and Computer Engineering from the National Technical University of Athens, Greece, in 2025, where his diploma thesis focused on Neuro-Symbolic AI for Visual Question Answering. He has contributed to several European research projects in the areas of machine learning and computer vision. He was an ML Researcher at Tech Hive Labs. He is currently a Machine Learning Engineer at Plaixus and serves under the Hellenic National Defence General Staff, Athens, Greece. His research interests include interpretable machine learning, end-to-end development and deployment of AI systems, and AI applications in medical, industrial, and security domains.
Table 17
Figure 18Figure 19Figure 20Figure 21
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Layer
Kernel
Filters
Output
Input
—
—
(T,1,128)
Rescaling
—
—
(T,1,128)
video_conv_0 (block)
3×3
⌊64α⌋
(T,1,⌊64α⌋)
video_conv_1 (block)
3×3
⌊96α⌋
(T,1,⌊96α⌋)
video_conv_2 (block)
3×3
⌊128α⌋
(T,1,⌊128α⌋)
Global avg. pool
—
—
(1,1,⌊128α⌋)
Appendix
TABLE XIII: Temporal-video encoder layers. Filter counts use the width multiplier α (1.0 on GRID, 0.5 on NAVIR). The output dimension D is 256 on GRID and 128 on NAVIR. On NAVIR, conv blocks are conv → batch normalisation → ReLU. On GRID, batch normalisation is omitted, leaving conv → ReLU.
Deep learning has greatly advanced automatic speech recognition (ASR), enabling widespread deployment on edge devices such as smartphones and smart home systems. However, the computational and energy demands of deep neural networks pose significant challenges for such resource-constrained deployments, introducing latency and limiting real-time interaction. Neuromorphic computing offers a promising solution by introducing activation sparsity through spiking neural networks (SNNs) and event-driven neural networks, converting dense operations into sparse computations. However, a study that evaluates the hardware benefits of different neuromorphic strategies remains lacking for ASR. This paper explores spiking and event-driven neuromorphic neural networks to improve activation sparsity in the state-of-the-art SpeechMamba model for ASR. We introduce an event-driven SpeechMamba with FATReLU activation, achieving over 60% activation sparsity with less than 1% accuracy degradation on LibriSpeech. We also propose a spiking SpeechMamba that attains over 70% sparsity while using 30% fewer parameters than comparable SNNs. Finally, we develop a cycle-accurate event-driven simulator enabling flexible algorithm-hardware co-exploration, which helps us identify computational bottlenecks and yields over 10% additional efficiency improvements.
Tauseef Ahmed, Tao Sun, Jeronimo Castrillon +2
Department of Advanced Computing Sciences, Maastricht University, Netherlands · Hardware-Efficient AI Team, IMEC, Netherlands · Chair for Compiler Construction, TU Dresden, Germany +1
Audio-Visual Speech Recognition takes two input modalities, acoustic and visual streams, where visual information from lip movements aids recognition when audio is noisy. Recently, LLM-based AVSR models have emerged as a promising paradigm by connecting pre-trained audio-visual encoders to an LLM, achieving strong results in clean conditions. However, these models are predominantly optimized for clean acoustic conditions, with limited attention to making the LLM backbone robust to noise. No explicit mechanism is employed to produce stable representations under corrupted audio, leading to performance degradation in noisy environments. To address this, we propose VIB-AVSR, which integrates Variational Information Bottleneck layers at targeted positions within the LLM backbone to regularize representations. VIB-AVSR reduces degradation under noisy conditions across multiple SNR levels and noise types, without requiring architectural modifications or additional training data.
We introduce LRS-VoxMM, an in-the-wild benchmark for audio-visual speech recognition (AVSR). The benchmark is derived from VoxMM, a dataset of diverse real-world spoken conversations with human-annotated transcriptions. We select AVSR-suitable samples and preprocess them in an LRS-style format for direct use in existing AVSR pipelines. Compared with commonly used benchmarks, LRS-VoxMM covers a more diverse range of scenarios and acoustic conditions. We also release distorted evaluation sets with additive noise, reverberation, and bandwidth limitation to support evaluation under severe acoustic degradation. Experimental results show that LRS-VoxMM is considerably harder than LRS3 and that the contribution of visual information becomes more evident as the audio signal degrades. LRS-VoxMM supports more realistic AVSR benchmarking and encourages further research on the role of visual information in challenging real-world conditions.
Doyeop Kwak, Jeongsoo Choi, Suyeon Lee +1
Korea Advanced Institute of Science and Technology, South Korea