Organizations: School of CSIT, Adelaide University, Adelaide, SA 5000, Australia · CAIAA, University of North Texas, Denton, TX 76203, USA · School of EECMS, Curtin University, Bentley, WA 6102, Australia · ACFR, The University of Sydney, Camperdown, NSW 2050, Australia
RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contribution magnitudes, while making both vulnerable to teacher judgment errors and preference variance, as supported by our theoretical analysis. To separate credit direction from its contribution magnitude, we introduce \textit{Decoupled Credit Self-Distillation (DCSD)}, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision. Specifically, we design belief-margin probing to determine credit direction and marginal information gain to quantify credit magnitude, enabling step-to-token credit assignment for policy optimization. Across 11 benchmarks, DCSD achieves the best overall scores against GRPO, OPSD, RLSD, and RLCSD. Compared with base models, DCSD improves the overall score by 8.45 points on mathematical reasoning and 7.01 points on multimodal reasoning, while correcting the credit direction for 6% of tokens and yielding a 1.5× reduction in token credit magnitude.
Figures & tables
Figure 1 : Overview of OPSD vs. DCSD credit. Direct OPSD can deviate from oracle credit in both direction and magnitude. DCSD decouples these roles to calibrate credit direction and contribution magnitude, reducing mismatch while preserving useful local credit.
Figure 2 : DCSD workflow. Given a student rollout, DCSD identifies reasoning steps from representation transitions, determines credit direction from answer-belief changes and magnitude from marginal information gain, and uses calibrated teacher signals to distribute step credit across tokens.
Method
Benchmarks
Overall
(a) Mathematical reasoning (Qwen3-4B)
AIME24
AIME25
AIME26
AMC23
MATH500
HMMT
Vanilla
65.83
54.17
52.50
92.50
88.10
26.67
81.40
GRPO
67.50
55.83
50.00
90.00
92.20
33.33
84.70
OPSD
50.00
46.67
45.83
87.50
88.60
26.67
80.11
RLSD
63.33
52.50
55.83
92.50
91.00
33.33
83.86
Table 1: Reasoning performance (mean@4). Best results are bold .
Deviation
Method
C0
C1
C2
C3
C4
Avg.
Top-1 disagreement (TDR, %)
OPSD
12.07
12.16
12.15
12.18
12.07
12.12
RLSD
9.69
10.31
10.26
10.31
10.22
10.16
DCSD ( Ours )
9.45
9.95
9.89
9.94
9.88
9.82
Support discrepancy (MAE, 10−3 )
OPSD
215.96
217.78
217.80
218.59
219.21
217.87
RLSD
186.15
200.17
199.76
200.04
206.19
198.46
DCSD ( Ours )
183.34
193.95
194.00
193.89
199.77
192.99
Table 2 : Token-level teacher-signal deviation from the condition-matched Qwen3-32B reference. Avg. averages C0–C4. Lower is better; best results are bold .
AIME24
AIME25
AIME26
Method
Pass@1
Pass@4
Length
Pass@1
Pass@4
Length
Pass@1
Pass@4
Length
RLCSD
63.33
73.33
11,450
50.00
70.00
12,946
60.00
70.00
12,465
DCSD ( Ours )
70.00
76.67
10,627
56.67
73.33
12,052
56.67
76.67
11,847
Table 3 : Reasoning accuracy (%) and efficiency (mean token length) across six benchmarks. Higher accuracy and shorter responses are better; best results are bold .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Math
VL
Base model
Qwen3-4B
Qwen3-VL-8B-Instruct
Training corpus
DAPO-17K
MMFineReason-123K
Prompt batch / rollouts per prompt
256 / 4
256 / 8
Prompt / response token limits
2,048 / 16,384
4,096 / 4,096
Rollout temperature / top- p / top- k
1.0/1.0 / off
1.0/1.0 / off
Validation temperature / top- p / top- k
0.7/0.95/20
0.7/0.8/20
Appendix
Table 4 : Training and validation settings for mathematical and vision–language reasoning.
Condition
Evidence
C0
Original question, without the additional user-message wrapper.
C1
The correct final answer is {g}.
C2
The correct final answer is {w}.
C3
The verified final result for this problem is {g}.
C4
A worked solution, followed by a blank line and the C1 sentence.
Appendix
Table 5 : Privileged evidence supplied under the C0–C4 diagnostic conditions. Here g is the gold answer and w is a deterministically constructed incorrect answer.
Baseline
MAE reduction
TDR reduction
OPSD
11.42%[9.84,13.20]
18.99%[17.65,20.48]
RLSD
2.76%[2.51,3.02]
3.30%[3.01,3.58]
Appendix
Table 6 : Relative reduction in reference-model error. Brackets contain paired-bootstrap 95% confidence intervals, expressed in percent.
Setting
GRPO
OPSD
RLSD
RLCSD
Base model / training data
Qwen3-4B (thinking) / DAPO-17K
Epochs / evaluated checkpoint
1 / step 70
Prompt batch / rollouts per prompt
256 / 4
Prompt / response limits (tokens)
2,048 / 16,384
Train temperature / top- p / top- k
1.0 / 1.0 / off
Validation T / top- p / top- k / n
0.7 / 0.95 / 20 / 1
Appendix
Table 7 : Training settings of the mathematical baselines implemented in this work. Top- k “off” corresponds to −1 . Training-time validation is not the final four-response evaluation.
Method
Generation
Old log- p
Actor
Other
Total
GRPO
0.8297
0.0760
0.3089
0.0008
1.2154
OPSD
0.9775
–
0.3603
<0.0001
1.3379
RLSD
0.8516
0.0661
0.3280
0.0027
1.2484
RLCSD ∗
0.8471
0.0823
0.0618
0.0923
1.0835
DCSD †
0.8505
0.0748
0.3765
0.3463
1.6481
Appendix
Table 8 : Logged training cost (GPU-hours per step) for mathematical reasoning.
State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China · Institute of Wireless Communications Technology, Shanghai Jiao Tong University, Shanghai, China