cs.NEOct 6, 2026

Common-Mode Errors Limit Low-Timestep Deep Spiking Q-Networks

Authors: Zijie Xu, Bingrui Guo, Yiding Sun, Yiting Dong, Zhile Yang, Zhaofei Yu

Organizations: Peking University Beijing, 100871, China

Abstract

Spiking neural networks (SNNs) offer sparse and event-driven computation, making them attractive for energy-constrained reinforcement learning (RL) on edge devices. In value-based RL, deep spiking Q-networks (DSQNs) combine such efficiency with action-value estimation for decision making. However, existing DSQNs often require multiple simulation timesteps for competitive performance, increasing computational and energy costs, whereas reducing the timesteps can cause substantial performance degradation. We investigate this degradation from the perspective of Q-value estimation errors. By decomposing errors across actions into common-mode and differential-mode components, we find that low-timestep DSQNs suffer disproportionately from common-mode errors shared across action values, which are particularly detrimental to temporal-difference learning through bootstrapped targets. Based on this finding, we propose Common-Mode Compensation Deep Spiking Q-Network (CMC-DSQN), which uses an auxiliary ANN to compensate for common-mode errors in the SNN outputs. At inference, greedy action selection can be performed directly from the SNN outputs, allowing the auxiliary ANN to be completely removed and preserving the energy efficiency of SNNs. Extensive experiments on Atari and MiniAtar environments demonstrate substantial performance improvements under low-timestep settings. CMC-DSQN outperforms state-of-the-art DSQN baselines by nearly 20%20\% at T=2T=2 and further surpasses the ANN baseline at T=4T=4.

Figures & tables

Explore similar work

CardsList
  1. Latency and accuracy tradeoffs in Spiking Neural Networks

    Sep 28, 2026Zhanglu Yan, Zixuan Zhu, Kaiwen Tang +3Spiking Neural NetworksTime-To-First-Spike

  2. Scalable Training of Continuous-Time Spiking Neural Networks with Differentiable Spike-Time Discretization

    Jul 16, 2026Yusuke Sakemi, Tomoya Takeuchi, Takeo Hosomi +1Spiking Neural NetworksTime-To-First-Spike