Organizations: Department of computer science, National University of Singapore · Shanghai Advanced Research Institute, Chinese Academy of Science · University of Chinese Academy of Sciences, Beijing · School of Artificial Intelligence, Shandong University
Spiking neural networks are attractive for low-power speech command recognition, yet their latency has received far less attention than their energy efficiency, and their multi-timestep execution is widely assumed to make them slower than quantized neural networks. This paper challenges the assumption that more local timesteps necessarily imply higher network latency. By overlapping computation across adjacent layers at the timestep level, SNNs may complete execution in less time than comparable bit-serial QNNs. However, this overlap relies on spikes firing on incomplete inputs, and a spike once generated cannot be withdrawn, so its error persists and reduces accuracy. Waiting for more input before firing would seem to improve accuracy at the cost of reduced overlap. Yet we find and prove that this intuition fails at some layers, where even a small increase in waiting can change spike timing and downstream computation, making the network both slower and less accurate. We therefore propose a Pipeline Delay Search method which selects each layer's delay by balancing task-level accuracy gains against added network latency. We then adapt the selected configurations through spike-based quantization-aware training and bounded tuning of firing thresholds and initial membrane potentials. Together, these steps form Falcon, a framework for Fine-grained Analysis of Latency and Controlled firing which systematically analyzes and optimizes SNN latency under a spatial analog compute-in-memory mapping with shared digital engines. We evaluate Falcon on GSCV2 and SSC, achieving competitive accuracies of 96.31 and 83.02 at modeled network-core latencies of 119.64 and 124.00us, respectively. Together, our analysis and results show that SNNs can compute more yet finish faster, and wait longer yet predict worse, highlighting why Falcon matters for both latency and accuracy.
Figures & tables
Figure 1: Cross-timestep pipelining and firing delay in SNNs. (a) Different layers process different timesteps in parallel, allowing earlier completion. (b) Delay-controlled IF kernel with δ sets how many input slots are accumulated before the first firing decision.
Figure 2: Effects of firing delay on Falcon-Large accuracy and latency.(a) Accuracy and latency under uniform delays. (b) Layerwise accuracy changes for delay increases 1→2 and 1→3 . (c) Three representative layer responses. (d) PDS candidate screening, with delay increases passing both score and accuracy-gain gates highlighted in green.
Figure 3: Delay search and training of Falcon-Large on GSC. (a) Validation accuracy and latency of Fastest, Balanced, and Accurate before SNN training. (b) Comparison with ten random delay schedules at the same core latency before SNN training. (c) Validation accuracy at each SNN training stage.
SSC Falcon-Medium
GSC Falcon-Medium
GSC Falcon-Large
Configuration
Acc.(%)
latency( μ s)
Acc.(%)
latency( μ s)
Acc.(%)
latency( μ s)
Matched QNN
83.97
131.46
96.68
130.86
96.92
202.98
Full-lookahead
83.94
253.04
96.68
252.44
96.91
385.44
PDS–Fastest
81.24
116.42
96.11
115.82
95.46
180.42
PDS–Balanced
83.02
124.00
96.31
119.64
96.19
184.24
PDS–Accurate
83.02
129.70
96.39
129.04
96.52
201.16
Table 1: Test accuracy and modeled network-core latency.
Method
Params. (M)
Accuracy (%)
Latency ( μ s)
T-BSO Liang et al. (2025) , Spiking VGG-11
9.23
96.12
1714.88
SpikeSCR Wang et al. (2024a) , 1L-16-256 (short)
1.63
94.71
278.47
SpikeSCR, 1L-16-256
1.63
95.90
468.11
SpikCommander Wang et al. (2026) , 1L-16-256
1.12
96.71
487.97
SpikCommander, 2L-16-256 (short)
2.13
96.27
645.63
SpikCommander, 2L-16-256
2.13
96.92
940.70
Table 2: Comparison on GSC in terms of model size, accuracy, and latency. As most prior works do not report hardware latency, we reproduce the latency of several representative state-of-the-art SNN models with digital arrays of the same size of Falcon. Detailed implementations and latency reconstruction are provided in Appendix F . For Falcon, results are reported as Non-pipelined / Balanced .
SSC
GSC
Method / Configuration
Params. (M)
Accuracy (%)
Params. (M)
Accuracy (%)
DCLS-Delays, 2L-2KC ( Hammouamri et al., 2024 )
1.40
80.16±0.09
1.40
95.00±0.06
DCLS-Delays, 3L-2KC
2.50
80.69±0.21
2.50
95.35±0.04
d-cAdLIF ( Deckers et al., 2024 )
0.70
80.23±0.07
0.61
95.69±0.03
SNN-Delays+ TR/NAR ( Zhang et al., 2024 )
2.50
81.02
2.50
95.62
CADAD, 3L ( Bai et al., 2026 )
0.60
80.69±0.24
0.60
95.58±0.15
Table 3: Comparison with prior SNN speech-recognition methods. Falcon reports Full-lookahead (direct conversion) / PDS–Accurate .“–” denotes an unreported or unverified value.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Effects of firing delay on Falcon-Medium accuracy and latency (GSC).(a) Accuracy and latency under uniform delays. (b) Layerwise accuracy changes for delay increases 1→2 and 1→3 . (c) Three representative layer responses. (d) PDS candidate screening, with delay increases passing both score and accuracy-gain gates highlighted in green.
Figure 5: Effects of firing delay on Falcon-Medium accuracy and latency (SSC).(a) Accuracy and latency under uniform delays. (b) Layerwise accuracy changes for delay increases 1→2 and 1→3 . (c) Three representative layer responses. (d) PDS candidate screening, with delay increases passing both score and accuracy-gain gates highlighted in green.
Figure 6: Delay search and training of Falcon-Medium on GSC. (a) Validation accuracy and latency of Fastest, Balanced, and Accurate before SNN training. (b) Comparison with ten random delay schedules at the same core latency before SNN training. (c) Validation accuracy at each SNN training stage.
Figure 7: Delay search and training of Falcon-Medium on SSC. (a) Validation accuracy and latency of Fastest, Balanced, and Accurate before SNN training. (b) Comparison with ten random delay schedules at the same core latency before SNN training. (c) Validation accuracy at each SNN training stage.
Figure 8: Validation accuracy during 30 epochs of QAT for PDS-selected Balanced schedules: (a) GSC-Medium, (b) GSC-Large, and (c) SSC-Medium. Dashed lines mark the ten-epoch budget, and highlighted points indicate the highest validation accuracy within 30 epochs.
Schedule
GSC Falcon-Medium
GSC Falcon-Large
SSC Falcon-Medium
PDS
96.31
96.19
83.02
Random 1
96.08
96.18
82.08
Random 2
96.33
96.02
82.39
Random 3
96.25
95.78
82.41
Random 4
96.24
95.81
81.66
Random 5
96.13
96.00
82.01
Appendix
Table 4: Test accuracy (%) after full adaptation with 30 epochs of QAT, 15 epochs of neuron-parameter tuning, and 10 epochs of joint training. Each random schedule matches the delay distribution and modeled core latency of the corresponding PDS Balanced schedule. All runs use the same training seed, and checkpoints are selected by validation accuracy.
Figure 9: Equal-budget QAT comparison on SSC with FALCON-Medium. The PDS-selected Balanced schedule and five random schedules share the same delay distribution and modeled core latency. (a) Validation accuracy over all ten training epochs. (b) A closer view of epochs 4–10.
Figure 10: Equal-budget QAT comparison on GSC with FALCON-Medium. The PDS-selected Balanced schedule and five random schedules share the same delay distribution and modeled core latency. (a) Validation accuracy over all ten training epochs. (b) A closer view of epochs 4–10.
Figure 11: Equal-budget QAT comparison on GSC with Falcon-Large. The PDS-selected Balanced schedule and five random schedules share the same delay distribution and modeled core latency. (a) Validation accuracy over all ten training epochs. (b) A closer view of epochs 4–10.
Figure 12: Analog energy breakdown.
Figure 13: Digital energy breakdown. Core energy includes array arithmetic, internal registers, control, and leakage during the evaluation windows.
Q/K prefix p
Core latency ( μ s)
Before val. (%)
After val. (%)
Before test (%)
After test (%)
7 (reference)
119.64
96.4332
–
96.3108
–
6
114.00
95.9523
96.2930
96.1654
96.3289
5
108.36
94.3192
96.1527
94.4480
96.1290
4
102.72
85.0215
95.8120
84.6524
95.9200
3
97.08
72.1872
91.6942
71.0404
91.1495
Appendix
Table 5: Q/K prefix ablation on GSC with Falcon-Medium.
Component
Falcon-Medium
Falcon-Large
Input-encoder convolutions
1
1
Pipelined-stem convolutions
2
2
Pipelined-stem FC layers
2
2
Transformer blocks
3
5
Hidden dimension
160
160
Attention heads
5
5
Appendix
Table 6: Architecture summary of the two Falcon variants. FC counts include the separate Q/K/V and output projections in every attention module, both feed-forward layers, the two stem FC layers, and the classifier.
Model
Schedule
Boundaries with δ=2
Boundaries with δ=3
GSC-Medium
Fastest
—
—
Balanced
Block 0: FFN1; Block 2: AttnOut
Block 2: Residual2
Accurate
Stem Conv2
Blocks 0 and 1: FFN1; Block 2: AttnOut and Residual2
GSC-Large
Fastest
—
—
Balanced
Block 2: Residual1; Block 3: FFN1
Block 4: Residual2
Accurate
Blocks 0, 1, 3, and 4: FFN1; Block 2: Residual1; Blocks 3 and 4: AttnOut
Stem Conv2; Block 2: FFN1; Block 4: Residual2
Appendix
Table 7: Layer-wise firing delays. Only searchable boundaries with δ=2 or δ=3 are listed; all other searchable boundaries use δ=1 . A dash indicates that no searchable boundary uses the corresponding delay. Transformer blocks are zero-indexed.
Figure 14: Latency Example of Falcon-Medium with balanced schedule
Spiking neural networks (SNNs) offer a biologically inspired computing paradigm with significant potential for energy-efficient neural processing. Among neural coding schemes of SNNs, Time-To-First-Spike (TTFS) coding, which encodes information through the precise timing of a neuron's first spike, provides exceptional activity sparsity and energy efficiency. However, existing TTFS models lack efficient training methods, suffering from high inference latency and limited performance, limiting their practicality on neuromorphic hardware. In this work, we propose latency coding, an extension of TTFS coding, and present a compatible framework that enables the efficient training of deep latency-coded SNNs by leveraging backpropagation through time (BPTT) algorithm. The framework includes: (1) a latency encoding (LE) module with feature extraction and straight-through estimators to address severe information loss in direct intensity-to-latency mapping; (2) relaxation of the strict single-spike constraint in intermediate layers to improve information propagation and gradient flow; and (3) a temporal adaptive decision (TAD) loss function that dynamically weights supervision signals based on the model's confidence, balancing the trade-off between speed and accuracy. Experimental results demonstrate that our method achieves competitive or superior accuracy compared with existing TTFS-coded SNNs with ultra-low inference latency and high energy efficiency. Latency-coded SNNs also demonstrate improved robustness against input perturbations. These findings highlight latency coding as a practical and hardware-friendly approach for fast and energy-efficient neuromorphic processing.
Yi Lu, Jianhao Ding, Zhaofei Yu
School of Software and Microelectronics, Peking University, Beijing, China · School of Computer Science, Peking University, Beijing, China · Institute for Artificial Intelligence and the School of Computer Science, Peking University, Beijing, China
Spiking Neural Networks (SNNs) are widely regarded as an energy-efficient paradigm for modeling and processing temporal and event-driven information. Incorporating delays in SNNs has been proven to be an effective mechanism for improving spike alignment in event-driven tasks. However, existing delay learning approaches predominantly assign static delays to individual synapses, resulting in a large number of delay parameters and limited adaptability to input-dependent activity dynamics. To this end, we propose a Congestion-Aware Dynamic Axonal Delay (CADAD) mechanism, which decomposes the delay into a channel-wise static base delay for temporal structuring and a global, activity-conditioned shift that dynamically regulates the state update rate under varying spike intensities. The delay parameters are learned using differentiable linear interpolation and discretized at inference time, preserving the benefits of dynamic delay modulation while incurring only minimal additional cost. Experiments on speech benchmarks, including the Spiking Heidelberg Dataset, Spiking Speech Commands, and Google Speech Commands, demonstrate that introducing congestion-aware delays into synaptic signal transmission effectively improves accuracy on temporal tasks, notably achieving 93.75% accuracy on SHD, 80.69% accuracy on SSC, and 95.58% on GSC-35, while reducing the parameter count by approximately 50% compared to state-of-the-art delay-based methods with the same architecture.
Dewei Bai, Hongxiang Peng, Yunyun Zeng +2
University of Electronic Science and Technology of China
Spiking neural networks (SNNs) exploit event-driven and addition-only computation to substantially improve efficiency for intelligent computation. A key temporal property of SNNs, elastic inference, allows outputs to emerge progressively, enabling responses to salient inputs much earlier than full evaluation. However, existing SNN-specific accelerators cannot capitalize on this property. Layer-by-layer designs emit outputs only after all layers are complete, while time-step-by-time-step designs rely on coarse-grained, layer-wise pipelines that require synchronizing all spines/tokens within a layer. This barrier prevents results from being forwarded immediately, delaying the earliest possible response and forfeiting the benefits of elastic inference. To address these challenges, we propose ELSA, a near-SRAM dataflow architecture that realizes true elastic inference through a fine-grained spine/token-wise pipeline and hardware optimizations tailored to SNNs. ELSA forwards each spine/token immediately upon production, forming a continuous streaming pipeline that substantially reduces the latency to the first response. To enhance this lightweight execution, ELSA introduces a bundled address event representation protocol to lower communication traffic of network-on-chip (NoC), and leverages mini-batch spiking Gustavson-product to cut memory access and exploit inherent sparsity. Combined with mapping and scheduling optimizations, ELSA achieves efficient, event-driven computation without compromising accuracy. Experiments show that SNNs can outperform quantized artificial neural networks (QANNs) while maintaining on-par accuracy. For a 4-bit ResNet-50, ELSA achieves 3.4× speedup and 13.6× higher energy efficiency over the SOTA QANN accelerator (ANT), and 2.9× speedup and 22.1× energy efficiency gains over the SOTA SNN accelerator (PAICORE).
Kang You, Chen Nie, Lee Jun Yan +6
Intelligent Computing Research Group, School of Computer Science, Shanghai Jiao Tong University, Shanghai, CN · Shanghai AI Laboratory, Shanghai, CN · School of Computer Science, Shanghai Jiao Tong University, Shanghai, CN +1