Organizations: School of CSIT, Adelaide University, Adelaide, SA 5000, Australia · CAIAA, University of North Texas, Denton, TX 76203, USA · School of EECMS, Curtin University, Bentley, WA 6102, Australia · ACFR, The University of Sydney, Camperdown, NSW 2050, Australia
RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contribution magnitudes, while making both vulnerable to teacher judgment errors and preference variance, as supported by our theoretical analysis. To separate credit direction from its contribution magnitude, we introduce \textit{Decoupled Credit Self-Distillation (DCSD)}, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision. Specifically, we design belief-margin probing to determine credit direction and marginal information gain to quantify credit magnitude, enabling step-to-token credit assignment for policy optimization. Across 11 benchmarks, DCSD achieves the best overall scores against GRPO, OPSD, RLSD, and RLCSD. Compared with base models, DCSD improves the overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning, while correcting the credit direction for 6% of tokens and yielding a 1.5× reduction in token credit magnitude.
Figures & tables
Figure 1 : Overview of OPSD vs. DCSD credit. Direct OPSD can deviate from oracle credit in both direction and magnitude. DCSD decouples these roles to calibrate credit direction and contribution magnitude, reducing mismatch while preserving useful local credit.
Figure 2 : DCSD workflow. Given a student rollout, DCSD identifies reasoning steps from representation transitions, determines credit direction from answer-belief changes and magnitude from marginal information gain, and uses calibrated teacher signals to distribute step credit across tokens.
Method
Benchmarks
Overall
(a) Mathematical reasoning (Qwen3-4B)
AIME24
AIME25
AIME26
AMC23
MATH500
HMMT
Vanilla
65.83
54.17
52.50
92.50
88.10
26.67
81.40
GRPO
67.50
55.83
50.00
90.00
92.20
33.33
84.70
OPSD
50.00
46.67
45.83
87.50
88.60
26.67
80.11
RLSD
63.33
52.50
55.83
92.50
91.00
33.33
83.86
Table 1: Reasoning performance (mean@4). Best results are bold .
Deviation
Method
C0
C1
C2
C3
C4
Avg.
Top-1 disagreement (TDR, %)
OPSD
12.07
12.16
12.15
12.18
12.07
12.12
RLSD
9.69
10.31
10.26
10.31
10.22
10.16
DCSD ( Ours )
9.45
9.95
9.89
9.94
9.88
9.82
Support discrepancy (MAE, 10−3 )
OPSD
215.96
217.78
217.80
218.59
219.21
217.87
RLSD
186.15
200.17
199.76
200.04
206.19
198.46
DCSD ( Ours )
183.34
193.95
194.00
193.89
199.77
192.99
Table 2 : Token-level teacher-signal deviation from the condition-matched Qwen3-32B reference. Avg. averages C0–C4. Lower is better; best results are bold .
AIME24
AIME25
AIME26
Method
Pass@1
Pass@4
Length
Pass@1
Pass@4
Length
Pass@1
Pass@4
Length
RLCSD
63.33
73.33
11,450
50.00
70.00
12,946
60.00
70.00
12,465
DCSD ( Ours )
70.00
76.67
10,627
56.67
73.33
12,052
56.67
76.67
11,847
Table 3 : Reasoning accuracy (%) and efficiency (mean token length) across six benchmarks. Higher accuracy and shorter responses are better; best results are bold .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Math
VL
Base model
Qwen3-4B
Qwen3-VL-8B-Instruct
Training corpus
DAPO-17K
MMFineReason-123K
Prompt batch / rollouts per prompt
256 / 4
256 / 8
Prompt / response token limits
2,048 / 16,384
4,096 / 4,096
Rollout temperature / top- p / top- k
1.0/1.0 / off
1.0/1.0 / off
Validation temperature / top- p / top- k
0.7/0.95/20
0.7/0.8/20
Appendix
Table 4 : Training and validation settings for mathematical and vision–language reasoning.
Condition
Evidence
C0
Original question, without the additional user-message wrapper.
C1
The correct final answer is {g}.
C2
The correct final answer is {w}.
C3
The verified final result for this problem is {g}.
C4
A worked solution, followed by a blank line and the C1 sentence.
Appendix
Table 5 : Privileged evidence supplied under the C0–C4 diagnostic conditions. Here g is the gold answer and w is a deterministically constructed incorrect answer.
Baseline
MAE reduction
TDR reduction
OPSD
11.42%[9.84,13.20]
18.99%[17.65,20.48]
RLSD
2.76%[2.51,3.02]
3.30%[3.01,3.58]
Appendix
Table 6 : Relative reduction in reference-model error. Brackets contain paired-bootstrap 95% confidence intervals, expressed in percent.
Setting
GRPO
OPSD
RLSD
RLCSD
Base model / training data
Qwen3-4B (thinking) / DAPO-17K
Epochs / evaluated checkpoint
1 / step 70
Prompt batch / rollouts per prompt
256 / 4
Prompt / response limits (tokens)
2,048 / 16,384
Train temperature / top- p / top- k
1.0 / 1.0 / off
Validation T / top- p / top- k / n
0.7 / 0.95 / 20 / 1
Appendix
Table 7 : Training settings of the mathematical baselines implemented in this work. Top- k “off” corresponds to −1 . Training-time validation is not the final four-response evaluation.
Method
Generation
Old log- p
Actor
Other
Total
GRPO
0.8297
0.0760
0.3089
0.0008
1.2154
OPSD
0.9775
–
0.3603
<0.0001
1.3379
RLSD
0.8516
0.0661
0.3280
0.0027
1.2484
RLCSD ∗
0.8471
0.0823
0.0618
0.0923
1.0835
DCSD †
0.8505
0.0748
0.3765
0.3463
1.6481
Appendix
Table 8 : Logged training cost (GPU-hours per step) for mathematical reasoning.
Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents, yet its trajectory-level credit assignment conflates heterogeneous per-turn outcomes into a single reward signal. On-policy distillation provides dense per-token supervision but is either teacher-bounded or prone to gradient concentration collapse. We introduce CrEST, a hierarchical credit assignment framework that retains RL's verifier-bounded ceiling while incorporating dense token-level signals from a privileged self-teacher. CrEST resolves credit at two levels: turn-segmented verified advantages address inter-turn dilution, while entropy-gated self-teacher modulation refines intra-turn token contributions. Experiments on BFCL V3 and WildToolBench show that CrEST consistently outperforms both RL and distillation baselines across two model scales, with the largest gains on long-trajectory and strict session-level metrics. Our work demonstrates that the teacher's role in policy optimization can be reduced from determining update directions to modulating update magnitudes, unlocking dense credit assignment without sacrificing the verifier-bounded ceiling.
Zechuan Wang, Siyuan Lu, Hongxuan Zhang +3
Zhejiang University · AWorld Team, Inclusion AI · Shanghai Innovation Institute +2
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in training LLMs for reasoning tasks, but representative methods such as GRPO assign uniform credit across all tokens, wasting gradient on routine tokens while under-crediting pivotal reasoning steps. Existing token-level credit assignment methods require resources beyond the model's own rollouts. GRPO variants rely on process reward models or ground-truth answers. Knowledge distillation assigns credit through per-token divergence but requires external teachers (On-Policy Distillation) or privileged information (On-Policy Self Distillation). However, these dependencies limit applicability in the pure RLVR setting. We observe that conditioning the model on its own verified trajectories induces a measurable per-token KL divergence between the original and conditioned distributions, and prove that distilling from a self-teacher constructed by verified trajectories leads to infeasible weighted-average solutions when multiple verified trajectories exist. We propose SC-GRPO (Self-Conditioned GRPO), which uses KL divergence mentioned before as a multiplicative weight on GRPO gradients. Across five benchmarks spanning math, code, and agentic tasks, SC-GRPO consistently outperforms 8.1% over GRPO and 5.9% over DAPO with stronger OOD performance. Moreover, SC-GRPO achieves higher performance than OPD.
Yingyu Shan, Yuhang Guo, Zihao Cheng +7
Beijing Institute of Technology · Beihang University · Independent Researcher
Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language models. Critic-free methods such as GRPO convert response-level rewards into advantages and uniformly broadcast them across tokens, overlooking their unequal contributions to the final outcome. On-policy self-distillation (OPSD) instead provides dense distributional supervision by minimizing the forward KL divergence between an unprivileged policy and a privileged self-teacher, implicitly assuming that the resulting likelihood shifts encode reliable answer-aligned information. We test this premise by fixing each sampled trajectory and re-scoring it under two opposing outcome conditions, one asserting correctness and the other incorrectness. Most affected tokens shift in the same direction under both conditions, with few sign reversals and substantial overlap in the induced optimization signals. Large shifts also concentrate on highly substitutable surface-form tokens, whereas tokens carrying problem-specific reasoning content are less sensitive. These findings show that privileged shifts fail to provide reliable answer-aligned directions, while their magnitudes primarily reflect counterfactual sensitivity rather than token-level learning value. Based on these observations, we propose Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for highly sensitive tokens and renormalizes token-level advantages to preserve both the original credit budget and verifier-determined direction. On long-CoT mathematical reasoning benchmarks, CSCR consistently outperforms GRPO baseline with the same number of policy updates. Targeted ablations further corroborate our diagnosis: privilege-induced directions are unreliable, moderate downweighting is most effective, and stronger modulation destabilizes optimization.
Qiangqiang He, Zhongheng Wu, ZiJian Wang
State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China · Institute of Wireless Communications Technology, Shanghai Jiao Tong University, Shanghai, China