Organizations: School of Information and Intelligent Science, Donghua University, Shanghai, China · School of Engineering, Westlake University, Hangzhou, China · Department of Electrical and Electronic Engineering, Imperial College London, London, United Kingdom · School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China
Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and expose credit-relevant information over time, a role we formalize as Temporal Credit Carriers (TCCs) and that spiking neural networks (SNNs) naturally fulfill through graded membrane traces and event-driven spikes. Based on this hypothesis, we propose SpikeCredit, an SNN-based framework for RL with sparse rewards that first performs task-adaptive TCC selection and then closes the loop between a fast TCC-reading pathway, where self-motion feedback constraint uses local behavior-grounded cues to constrain transition-level credit recovery, and a slow TCC-writing pathway, where credit-targeted trace alignment feeds recovered credit back into the actor to make future TCC dynamics more credit-readable. Across sparse-reward MuJoCo tasks, SpikeCredit improves Last10 return over sparse SNN baselines by +1169% on Ant, +953% on Hopper, +723% on Swimmer, and +1781% on Walker2d, and exceeds the dense-reward baseline on Swimmer by +113%. Mechanistic analyses further show substantially stronger alignment with dense rewards than the sparse SNN baseline. These results position spiking dynamics as credit-preserving substrates for sparse-reward RL.
Figures & tables
Figure 1: Motivation and core idea of SpikeCredit. Inspired by biological fast-slow learning, SpikeCredit treats spiking actor dynamics as task-adaptive temporal credit carriers (TCCs): the fast pathway reads credit from current dynamics, while the slow pathway writes recovered targets to reshape future dynamics, forming a closed read-write loop.
Figure 2: Overview of SpikeCredit. Task statistics select membrane- or spike-based TCCs. The fast pathway uses Self-Motion Feedback Constraint (SMF) to read credit from the selected dynamics and generate reshaped rewards and replay targets, while the slow pathway uses Credit-Targeted Trace Alignment (CTT) to write these targets back into the actor, progressively improving future TCC readability.
r^t=RS^tTCC,+,yt=log(TS^tTCC,+).
Algorithm 1 Closed-Loop Optimization in SpikeCredit
Figure 3: Learning curves of SpikeCredit, sparse-reward SNN and ANN baselines on four MuJoCo tasks. Lines and shaded regions denote the mean and standard deviation over five seeds, respectively.
Figure 4: Overall performance on four MuJoCo tasks. (a) Last10 returns against the sparse SNN baseline. (b) Learning curves against the dense SNN upper bound. (c) Comparisons with three alternative credit-assignment methods. Mean ± standard deviation over five seeds.
Figure 5: Task-adaptive TCC selection. Left: task event scores determine the carrier. Right: Last10 returns verify that the selected candidate performs best on each task. Results are averaged over five seeds.
Figure 6: Alignment between credit signals and dense rewards on Ant-v4. Left: transition-wise correlations. Right: temporal profiles over an episode. Dense rewards are used only as a diagnostic reference.
Figure 7: Task-dependent SMF proxy attribution. Top: feature importance over st , Δst , and ∥at∥22 . Bottom: attribution aggregated by physical feature groups. Semantic groupings of observation fields used for self-motion proxy attribution are provided in the Appendix.
Methods
Last10 Return
Peak Return
Peak Step
SpikeCredit
1669.9±402.8
1777.5±466.8
995K
w/o st
544.4±389.7
738.1±91.5
20K
w/o Δst
1106.3±470.5
1212.2±490.5
985K
w/o ∥at∥22
845.8±672.7
939.5±654.9
980K
w/o SMF
192.7±78.6
738.1±91.5
20K
SpikeCredit
1669.9±402.8
1777.5±466.8
995K
Table 1: Ablation of SMF inputs and CTT on Ant-v4. Results are averaged over five seeds.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Rendered views of the four MuJoCo continuous-control environments used in our experiments: (a) Ant-v4, (b) Hopper-v4, (c) Swimmer-v4, and (d) Walker2d-v4.
Parameter
Value
α
0.5
Vth
0.5
w
0.5
θv
−0.172
θu
0.529
θr
0.021
Appendix
Table 2: Dynamic neuron parameters used in all SNN experiments.
Table 6: Task-specific mapping from Gymnasium MuJoCo-v4 observation dimensions to semantic proxy-feature categories.
Figure 9: Hyperparameter sensitivity of SpikeCredit on Ant-v4. We vary λalign , λsparse , and λCTT while keeping other settings fixed, and report Last10 evaluation returns over five seeds. Dashed lines denote Sparse SNN and Mean Redistribution baselines. SpikeCredit consistently outperforms both baselines across all tested values, indicating that its gains are not tied to a narrow hyperparameter choice.
In many modern applications of reinforcement learning (RL), the natural reward for a task of interest is inherently sparse: a reward of 0 is given everywhere except when the task is completed, when a reward of +1 is given. Training a policy to maximize such a sparse reward requires solving a challenging credit assignment problem, leading to slow or ineffective RL improvement. We propose a simple approach to transform a sparse outcome reward into a dense process reward. Our approach relies on training a discriminator to distinguish between previous successful and unsuccessful episodes, and using this discriminator to incentivize the RL-learned policy to match the state-action visitations of successful episodes, while avoiding those of unsuccessful episodes. By incentivizing the policy to match the visitations over all states, not just those that correspond to task success, this reward provides dense feedback on whether progress is being made towards task completion, and, we show, provably achieves this without changing the optimal policy. Focusing on finetuning of robotic control policies, we demonstrate that our approach leads to significantly faster RL finetuning performance on both simulated and real-world manipulation tasks, as compared to simply maximizing the sparse outcome reward.
The temporal lag between actions and their long-term consequences makes credit assignment a challenge when learning goal-directed behaviors from data. Generative world models capture the distribution of future states an agent may visit, indicating that they have captured temporal information. How can that temporal information be extracted to perform credit assignment? In this paper, we formalize how the temporal information stored in world models encodes the underlying geometry of the world. Leveraging optimal transport, we extract this geometry from a learned model of the occupancy measure into a reward function that captures goal-reaching information. Our resulting method, Occupancy Reward Shaping, largely mitigates the problem of credit assignment in sparse reward settings. ORS provably does not alter the optimal policy, yet empirically improves performance by 2.2x across 13 diverse long-horizon locomotion and manipulation tasks. Moreover, we demonstrate the effectiveness of ORS in the real world for controlling nuclear fusion on 3 Tokamak control tasks. Code: https://github.com/aravindvenu7/occupancy_reward_shaping; Website: https://aravindvenu7.github.io/website/ors/
Aravind Venugopal, Jiayu Chen, Xudong Wu +3
Carnegie Mellon University · The University of Hong Kong · INFIFORCE Intelligent Technology +1
Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult. Existing scalable RL methods typically assign trajectory-level rewards uniformly across tokens, while recent entropy-aware approaches either rely on coarse detached heuristics or directly optimize true entropy, which can introduce non-local gradient components misaligned with sampled-token policy updates. We propose Adaptive Credit Policy Optimization (ACPO), a token-level credit assignment framework based on a mode-local surrogate entropy. ACPO asymmetrically modulates policy updates by emphasizing uncertain decisions in successful rollouts and overconfident tokens in failed rollouts. We show that the surrogate admits deterministic entropy bounds and, under modal alignment and proximal updates, preserves the policy-gradient direction to leading order. Experiments on mathematical reasoning and coding benchmarks, including AIME 2025 and HumanEvalPro, show that ACPO consistently improves over strong RL baselines such as DAPO, GTPO, and SAPO.