Spiking neural networks (SNNs) offer sparse and event-driven computation, making them attractive for energy-constrained reinforcement learning (RL) on edge devices. In value-based RL, deep spiking Q-networks (DSQNs) combine such efficiency with action-value estimation for decision making. However, existing DSQNs often require multiple simulation timesteps for competitive performance, increasing computational and energy costs, whereas reducing the timesteps can cause substantial performance degradation. We investigate this degradation from the perspective of Q-value estimation errors. By decomposing errors across actions into common-mode and differential-mode components, we find that low-timestep DSQNs suffer disproportionately from common-mode errors shared across action values, which are particularly detrimental to temporal-difference learning through bootstrapped targets. Based on this finding, we propose Common-Mode Compensation Deep Spiking Q-Network (CMC-DSQN), which uses an auxiliary ANN to compensate for common-mode errors in the SNN outputs. At inference, greedy action selection can be performed directly from the SNN outputs, allowing the auxiliary ANN to be completely removed and preserving the energy efficiency of SNNs. Extensive experiments on Atari and MiniAtar environments demonstrate substantial performance improvements under low-timestep settings. CMC-DSQN outperforms state-of-the-art DSQN baselines by nearly 20% at T=2 and further surpasses the ANN baseline at T=4.
Figures & tables
Figure 1: Comparison of existing directly trained DSQNs on the MiniAtar environments ( Young and Tian, 2019 ) , including DSQN ( Liu et al., 2022 ) , pbLN ( Sun et al., 2022 ) , DATSQN ( Ghoreishee et al., 2026 ) , CaRe-BN ( Xu et al., 2026c ) , and TP-DSQN ( Chowdhury et al., 2022 ) . Performance is averaged across all environments and normalized by the ANN.
Figure 2: Analysis on deterministic CliffWalking. (a) Environment. (b,c) Greedy trajectories of the ANN and low-timestep SNN, respectively, aggregated over 20 random seeds. Yellow segments denote trajectories, and red crosses mark states leading to infinite loops. (d,e) State-wise common-mode error ratios of the ANN and SNN with respect to the optimal action-value function Q∗ .
Figure 3: Learning dynamics and Q-value estimation errors on CliffWalking environment. (a) Fraction of successful seeds at each evaluation point. (b) Cumulative fraction of seeds that have solved the task at least once. Curves in (a) and (b) are smoothed for visual clarity. (c) Common- and differential-mode RMSEs with respect to the exact Qπ of each learned deterministic greedy policy.
Figure 4: Learning dynamics and Q-value estimation errors on CliffWalking. (a) Fraction of successful seeds at each evaluation point. (b) Cumulative fraction of seeds that have solved the task at least once. Curves in (a) and (b) are smoothed for visual clarity. (c) Common- and differential-mode RMSEs with respect to the exact optimal Q-function Q∗ . Unlike Fig. 3 (c), which uses the exact Qπ of each learned policy as the reference, this panel evaluates all methods against Q∗ .
Timesteps
Method
BeamRider
Breakout
Enduro
Pong
Q*bert
Seaquest
SpaceInvaders
APR
–
Random
354.64
0.97
0.00
-20.26
162.08
60.20
127.83
0.00%
DQN(ANN)
6842.99
364.51
295.79
19.86
9224.17
1960.44
1492.50
100.00%
T=2
DSQN
4279.90
276.11
280.97
17.60
1300.21
2101.49
806.36
70.75%
CMC-DSQN
4136.44
382.90
316.87
18.34
4186.33
2358.57
1183.40
87.06%
T=3
DSQN
4705.58
292.41
273.83
17.23
4076.25
2863.83
935.96
83.32%
CMC-DSQN
5284.37
398.73
313.55
19.59
8782.83
1766.93
1401.10
95.57%
Table 1: Average maximum return over three random seeds and average performance ratio (APR) on Atari environments after 30M environment transitions.
Timesteps
methods
Asterix
Breakout
Freeway
Seaquest
SpaceInvaders
APR
—
Random
0.493±0.902
0.366±0.618
0.370±0.586
0.116±0.427
4.057±3.199
0.00%
DQN (ANN)
32.22±4.94
27.76±4.98
62.58±0.34
49.70±8.15
138.34±3.87
100.00%
T=2
DSQN-CVT
1.89±1.14
4.11±2.40
19.17±17.72
1.36±0.75
3.82±3.77
10.12%
DCGS
18.36±3.55
11.64±2.21
59.22±0.77
26.17±2.84
60.56±20.89
57.33%
CRPI
21.02±2.36
14.92±0.89
59.70±0.64
28.56±2.44
82.02±21.38
65.72%
DSQN
14.50±3.04
23.90±1.22
56.92±1.06
22.84±15.67
100.76±11.51
67.76%
Table 2: Average maximum return over five random seeds and average performance ratio (APR) on MiniAtar environments after 5M environment transitions, where ± denotes one standard deviation.
Figure 5: Average normalized return across MiniAtar environments. (a)–(c) Comparison of different DSQN variants at T=2,4,6 , respectively. (d) Comparison of the vanilla ANN DQN and its common-mode compensated variant. All curves are smoothed uniformly for visual clarity.
Spiking neural networks are attractive for low-power speech command recognition, yet their latency has received far less attention than their energy efficiency, and their multi-timestep execution is widely assumed to make them slower than quantized neural networks. This paper challenges the assumption that more local timesteps necessarily imply higher network latency. By overlapping computation across adjacent layers at the timestep level, SNNs may complete execution in less time than comparable bit-serial QNNs. However, this overlap relies on spikes firing on incomplete inputs, and a spike once generated cannot be withdrawn, so its error persists and reduces accuracy. Waiting for more input before firing would seem to improve accuracy at the cost of reduced overlap. Yet we find and prove that this intuition fails at some layers, where even a small increase in waiting can change spike timing and downstream computation, making the network both slower and less accurate. We therefore propose a Pipeline Delay Search method which selects each layer's delay by balancing task-level accuracy gains against added network latency. We then adapt the selected configurations through spike-based quantization-aware training and bounded tuning of firing thresholds and initial membrane potentials. Together, these steps form Falcon, a framework for Fine-grained Analysis of Latency and Controlled firing which systematically analyzes and optimizes SNN latency under a spatial analog compute-in-memory mapping with shared digital engines. We evaluate Falcon on GSCV2 and SSC, achieving competitive accuracies of 96.31 and 83.02 at modeled network-core latencies of 119.64 and 124.00us, respectively. Together, our analysis and results show that SNNs can compute more yet finish faster, and wait longer yet predict worse, highlighting why Falcon matters for both latency and accuracy.
Zhanglu Yan, Zixuan Zhu, Kaiwen Tang +3
Department of computer science, National University of Singapore · Shanghai Advanced Research Institute, Chinese Academy of Science · University of Chinese Academy of Sciences, Beijing +1
Continuous-time spiking neural networks (SNNs) provide an event-driven framework for temporal computation, computational neuroscience, and neuromorphic hardware. However, training deep continuous-time SNNs is severely constrained by the memory required for exact spike-time computation, which evaluates and retains candidate firing times over intervals determined by presynaptic spike ordering. Here we introduce a memory-efficient training framework based on differentiable spike-time discretization (DSTD) for leaky integrate-and-fire neurons with general membrane and synaptic time constants. DSTD maps irregular presynaptic spikes onto differentiable weighted events at fixed time points, replacing the input-dependent candidate dimension with M fixed time intervals while accurately approximating continuous-time membrane-potential dynamics. This reduces candidate-related activation memory from O(NoutNin) to O(NoutM) in the case of time-to-first-spike (TTFS) coding, where Nin and Nout denote the numbers of presynaptic and postsynaptic neurons, respectively. We further introduce synfire-chain-inspired temporal regularization that organizes layer-wise firing windows, mitigates dead-neuron failures, and enables pipeline-like processing. In dense LIF layers, DSTD reduced peak memory consumption by up to approximately 100-fold and training time by up to approximately 20-fold compared with exact spike-time computation. Together, these methods allowed us to train 9-layer convolutional SNNs on CIFAR-10 and 20-layer convolutional SNNs on Fashion-MNIST on a single GPU.
Yusuke Sakemi, Tomoya Takeuchi, Takeo Hosomi +1
Research Center for Mathematical Engineering, Chiba Institute of Technology, Narashino, Japan · International Research Center for Neurointelligence (WPI-IRCN), The University of Tokyo, Tokyo, Japan · NEC Corporation, Kawasaki, Japan
Binary spike coding enables sparse and event-driven computation in spiking neural networks (SNNs), yet its 1-bit-per-timestep representation fundamentally limits information throughput. This bottleneck becomes increasingly restrictive in deep architectures under short simulation horizons. We propose the Quantized Burst-LIF (QB-LIF) neuron, which reformulates burst spiking as a saturated uniform quantization of membrane potentials with a learnable scale. Instead of relying on predefined multi-threshold structures, QB-LIF treats the quantization scale as a trainable parameter, allowing each layer to autonomously adapt its spiking resolution to the underlying membrane-potential statistics. To preserve hardware efficiency, we introduce an absorbable scale strategy that folds the learned quantized scale into synaptic weights during inference, maintaining a strict accumulate-only (AC) execution paradigm. To enable stable optimization in the discrete multi-level space, we further design ReLSG-ET, a rectified-linear surrogate gradient with exponential tails that sustains gradient flow across burst intervals. Extensive experiments on static (CIFAR-10/100, ImageNet) and event-driven (CIFAR10-DVS, DVS128-Gesture) benchmarks demonstrate that QB-LIF consistently outperforms binary and fixed-burst SNNs, achieving higher accuracy under ultra-low latency while preserving neuromorphic compatibility.