cs.CLOct 7, 2026

SpikingVLA: Asynchronous Spiking Vision-Language-Action Models

Authors: Jingya Wang, Dehao Zhang, Shuai Wang, Malu Zhang, Yang Yang, Haizhou Li

Organizations: University of Electronic Science and Technology of China · Shenzhen Loop Area Institute · The Chinese University of Hong Kong (Shenzhen)

Abstract

ANN-to-SNN conversion offers a practical route toward energy-efficient spiking Vision-Language-Action (VLA) models by bypassing the substantial cost of training large-scale SNNs from scratch. However, existing methods often require many timesteps to maintain competitive performance, resulting in substantial inference latency for real-time VLA deployment. To address this challenge, we introduce SpikingVLA, an ANN-to-SNN conversion framework that enables accurate and low-latency spiking VLA inference. Specifically, we propose a Dendritic Integrate-and-Fire (DIF) neuron that alleviates channel-wise activation outliers through dendritic mixing and adaptive somatic firing, enabling accurate ANN-to-SNN conversion with fewer timesteps. Building on DIF neurons, we further introduce an asynchronous execution mechanism that overlaps temporal computation across VLA components, reducing synchronization overhead and latency. Extensive experiments demonstrate that SpikingVLA achieves competitive navigation performance with substantially improved inference efficiency. Compared with existing spiking VLA methods, SpikingVLA improves SR and SPL by 11.9% and 12.6%, respectively, while reducing first-action latency by 11.2×\times. These results establish SpikingVLA as a practical framework for deploying pretrained VLA models with high-performance and low-latency spiking inference.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks

    Jun 26, 2026Ruiqi Song, Dujun Nie, Siyu Teng +7Diffusion-Based Vision-Language-ActionsFast Inference

  2. VL2Spike: Spike-driven Distillation from VLMs for Low-Power Visual Perception in Embodied AI

    Jun 14, 2026Zinan Liu, Eric Zheng, Soumyaratna Debnath +3Spiking Neural NetworksTime-To-First-Spike

  3. Understanding Asynchronous Inference Methods for Vision-Language-Action Models

    May 4, 2026Ayoub AgouzoulInference LatencyFast Inference