Organizations: School of Information and Intelligent Science, Donghua University, Shanghai, China · School of Engineering, Westlake University, Hangzhou, China · Department of Electrical and Electronic Engineering, Imperial College London, London, United Kingdom · School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China
Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and expose credit-relevant information over time, a role we formalize as Temporal Credit Carriers (TCCs) and that spiking neural networks (SNNs) naturally fulfill through graded membrane traces and event-driven spikes. Based on this hypothesis, we propose SpikeCredit, an SNN-based framework for RL with sparse rewards that first performs task-adaptive TCC selection and then closes the loop between a fast TCC-reading pathway, where self-motion feedback constraint uses local behavior-grounded cues to constrain transition-level credit recovery, and a slow TCC-writing pathway, where credit-targeted trace alignment feeds recovered credit back into the actor to make future TCC dynamics more credit-readable. Across sparse-reward MuJoCo tasks, SpikeCredit improves Last10 return over sparse SNN baselines by +1169% on Ant, +953% on Hopper, +723% on Swimmer, and +1781% on Walker2d, and exceeds the dense-reward baseline on Swimmer by +113%. Mechanistic analyses further show substantially stronger alignment with dense rewards than the sparse SNN baseline. These results position spiking dynamics as credit-preserving substrates for sparse-reward RL.
Figures & tables
Figure 1: Motivation and core idea of SpikeCredit. Inspired by biological fast-slow learning, SpikeCredit treats spiking actor dynamics as task-adaptive temporal credit carriers (TCCs): the fast pathway reads credit from current dynamics, while the slow pathway writes recovered targets to reshape future dynamics, forming a closed read-write loop.
Figure 2: Overview of SpikeCredit. Task statistics select membrane- or spike-based TCCs. The fast pathway uses Self-Motion Feedback Constraint (SMF) to read credit from the selected dynamics and generate reshaped rewards and replay targets, while the slow pathway uses Credit-Targeted Trace Alignment (CTT) to write these targets back into the actor, progressively improving future TCC readability.
r^t=RS^tTCC,+,yt=log(TS^tTCC,+).
Algorithm 1 Closed-Loop Optimization in SpikeCredit
Figure 3: Learning curves of SpikeCredit, sparse-reward SNN and ANN baselines on four MuJoCo tasks. Lines and shaded regions denote the mean and standard deviation over five seeds, respectively.
Figure 4: Overall performance on four MuJoCo tasks. (a) Last10 returns against the sparse SNN baseline. (b) Learning curves against the dense SNN upper bound. (c) Comparisons with three alternative credit-assignment methods. Mean ± standard deviation over five seeds.
Figure 5: Task-adaptive TCC selection. Left: task event scores determine the carrier. Right: Last10 returns verify that the selected candidate performs best on each task. Results are averaged over five seeds.
Figure 6: Alignment between credit signals and dense rewards on Ant-v4. Left: transition-wise correlations. Right: temporal profiles over an episode. Dense rewards are used only as a diagnostic reference.
Figure 7: Task-dependent SMF proxy attribution. Top: feature importance over st , Δst , and ∥at∥22 . Bottom: attribution aggregated by physical feature groups. Semantic groupings of observation fields used for self-motion proxy attribution are provided in the Appendix.
Methods
Last10 Return
Peak Return
Peak Step
SpikeCredit
1669.9±402.8
1777.5±466.8
995K
w/o st
544.4±389.7
738.1±91.5
20K
w/o Δst
1106.3±470.5
1212.2±490.5
985K
w/o ∥at∥22
845.8±672.7
939.5±654.9
980K
w/o SMF
192.7±78.6
738.1±91.5
20K
SpikeCredit
1669.9±402.8
1777.5±466.8
995K
Table 1: Ablation of SMF inputs and CTT on Ant-v4. Results are averaged over five seeds.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Rendered views of the four MuJoCo continuous-control environments used in our experiments: (a) Ant-v4, (b) Hopper-v4, (c) Swimmer-v4, and (d) Walker2d-v4.
Parameter
Value
α
0.5
Vth
0.5
w
0.5
θv
−0.172
θu
0.529
θr
0.021
Appendix
Table 2: Dynamic neuron parameters used in all SNN experiments.
Table 6: Task-specific mapping from Gymnasium MuJoCo-v4 observation dimensions to semantic proxy-feature categories.
Figure 9: Hyperparameter sensitivity of SpikeCredit on Ant-v4. We vary λalign , λsparse , and λCTT while keeping other settings fixed, and report Last10 evaluation returns over five seeds. Dashed lines denote Sparse SNN and Mean Redistribution baselines. SpikeCredit consistently outperforms both baselines across all tested values, indicating that its gains are not tied to a narrow hyperparameter choice.