SpikingVLA: Asynchronous Spiking Vision-Language-Action Models
Organizations: University of Electronic Science and Technology of China · Shenzhen Loop Area Institute · The Chinese University of Hong Kong (Shenzhen)
Abstract
ANN-to-SNN conversion offers a practical route toward energy-efficient spiking Vision-Language-Action (VLA) models by bypassing the substantial cost of training large-scale SNNs from scratch. However, existing methods often require many timesteps to maintain competitive performance, resulting in substantial inference latency for real-time VLA deployment. To address this challenge, we introduce SpikingVLA, an ANN-to-SNN conversion framework that enables accurate and low-latency spiking VLA inference. Specifically, we propose a Dendritic Integrate-and-Fire (DIF) neuron that alleviates channel-wise activation outliers through dendritic mixing and adaptive somatic firing, enabling accurate ANN-to-SNN conversion with fewer timesteps. Building on DIF neurons, we further introduce an asynchronous execution mechanism that overlaps temporal computation across VLA components, reducing synchronization overhead and latency. Extensive experiments demonstrate that SpikingVLA achieves competitive navigation performance with substantially improved inference efficiency. Compared with existing spiking VLA methods, SpikingVLA improves SR and SPL by 11.9% and 12.6%, respectively, while reducing first-action latency by 11.2. These results establish SpikingVLA as a practical framework for deploying pretrained VLA models with high-performance and low-latency spiking inference.
Figures & tables
| Type | Method | TF | R2R Val-Unseen | RxR Val-Unseen | Total T | ||||||
| NE | OS | SR | SPL | NE | SR | SPL | nDTW | ||||
| ANN | StreamVLN ( Wei et al., 2025 ) | ✗ | 4.98 | 64.2 | 56.9 | 51.9 | 6.22 | 52.9 | 46.0 | 61.9 | - |
| InternVLA-N1 ( Wei et al., 2026 ) | ✗ | 4.05 | 70.7 | 64.3 | 58.5 | 4.58 | 61.4 | 51.8 | 70.0 | - | |
| ETPNav ( An et al., 2024 ) | ✗ | 4.71 | 65.0 | 57.0 | 49.0 | 5.64 | 54.7 | 44.8 | 61.9 | - | |
| HNR ( Wang et al., 2024 ) | ✗ | 4.42 | 67.0 | 61.0 | 51.0 | 5.50 | 56.3 | 46.7 | 63.5 | - | |
| NaVILA ( Cheng et al., 2024 ) | ✗ | 5.28 | 61.5 | 53.9 | 49.3 | 6.12 | 52.3 | 46.1 | 61.0 | - | |
| Type | Method | TF | Closed-Loop Navigation | Deployment Efficiency | |||||
| NE | OS | SR | SPL | Mem. | TTFA | Total T | |||
| ANN | NaVILA-Blind ( Cheng et al., 2024 ) | ✗ | 6.03 | 49.0 | 36.2 | 33.3 | 16126.2 | - | - |
| NaVILA-Vision ( Cheng et al., 2024 ) | ✗ | 5.49 | 58.7 | 50.2 | 45.5 | 16126.2 | - | - | |
| NaVILA-R ( Song et al., 2026 ) | ✗ | 6.29 | 52.1 | 36.5 | 29.5 | 16126.2 | - | - | |
| NaVILA-AWQ ( Cheng et al., 2024 ) | ✗ | 6.58 | 48.2 | 32.8 | 27.4 | 5846.9 | - | - | |
| SNN | SpikeVLA ( Song et al., 2026 ) | ✗ | 6.02 | 53.6 | 32.7 | 28.5 | 6251.5 | 186.1 | 500-600 |
| Module | Neuron | Number of Thresholds | |||
| Vision Encoder | MT | 1.13 | 0.99 | 0.74 | 0.52 |
| DIF | 1.15 | 0.74 | 0.39 | 0.11 | |
| LLaMA Decoder | MT | 0.48 | 0.40 | 0.31 | 0.24 |
| DIF | 0.52 | 0.36 | 0.22 | 0.11 | |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Vision Encoder | Projector | Language Decoder | Actor Network | Total |
| NaVILA-Vision (ANN) | 24.53 | 0.51 | 120.33 | 2.90 | 148.27 |
| SpikingVLA ( ) | 3.87 | 0.31 | 19.05 | 0.45 | 23.68 |
| SpikingVLA ( ) | 7.70 | 0.35 | 37.84 | 0.91 | 46.80 |
| Vision Encoder | LLaMA Decoder | |||||
| Uncalib. | MT-Calib. | DIF-Calib. | Uncalib. | MT-Calib. | DIF-Calib. | |
| 1 | 1.00 | 1.13 | 1.15 | 1.30 | 0.48 | 0.52 |
| 2 | 1.02 | 0.99 | 0.74 | 1.31 | 0.40 | 0.36 |
| 4 | 1.07 | 0.74 | 0.39 | 1.26 | 0.31 | 0.22 |
| 8 | 1.12 | 0.52 | 0.19 | 1.24 | 0.24 | 0.11 |