cs.LGSep 28, 2026

Latency and accuracy tradeoffs in Spiking Neural Networks

Authors: Zhanglu Yan, Zixuan Zhu, Kaiwen Tang, Yuyang Cai, Qianhui Liu, Weng-Fai Wong

Organizations: Department of computer science, National University of Singapore · Shanghai Advanced Research Institute, Chinese Academy of Science · University of Chinese Academy of Sciences, Beijing · School of Artificial Intelligence, Shandong University

Abstract

Spiking neural networks are attractive for low-power speech command recognition, yet their latency has received far less attention than their energy efficiency, and their multi-timestep execution is widely assumed to make them slower than quantized neural networks. This paper challenges the assumption that more local timesteps necessarily imply higher network latency. By overlapping computation across adjacent layers at the timestep level, SNNs may complete execution in less time than comparable bit-serial QNNs. However, this overlap relies on spikes firing on incomplete inputs, and a spike once generated cannot be withdrawn, so its error persists and reduces accuracy. Waiting for more input before firing would seem to improve accuracy at the cost of reduced overlap. Yet we find and prove that this intuition fails at some layers, where even a small increase in waiting can change spike timing and downstream computation, making the network both slower and less accurate. We therefore propose a Pipeline Delay Search method which selects each layer's delay by balancing task-level accuracy gains against added network latency. We then adapt the selected configurations through spike-based quantization-aware training and bounded tuning of firing thresholds and initial membrane potentials. Together, these steps form Falcon, a framework for Fine-grained Analysis of Latency and Controlled firing which systematically analyzes and optimizes SNN latency under a spatial analog compute-in-memory mapping with shared digital engines. We evaluate Falcon on GSCV2 and SSC, achieving competitive accuracies of 96.31 and 83.02 at modeled network-core latencies of 119.64 and 124.00us, respectively. Together, our analysis and results show that SNNs can compute more yet finish faster, and wait longer yet predict worse, highlighting why Falcon matters for both latency and accuracy.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Latency Coding for Efficient and Low-Latency Deep Spiking Neural Networks

    Mar 24, 2026Yi Lu, Jianhao Ding, Zhaofei YuSpiking Neural NetworksTime-To-First-Spike

  2. Congestion-Aware Dynamic Axonal Delay for Spiking Neural Networks

    May 2, 2026Dewei Bai, Hongxiang Peng, Yunyun Zeng +2Spiking Neural NetworksTime-To-First-Spike

  3. ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing

    May 20, 2026Kang You, Chen Nie, Lee Jun Yan +6Spiking Neural NetworksNeuromorphic Computing