Organizations: Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University · School of Artificial Intelligence, Beihang University · Tsinghua University
Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 % and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD
Figures & tables
Figure 1: Decision–Timestamp Mismatch. Functional decisions need not share timestamps across different sibling rollouts and can span multiple turns within individual trajectories.
Figure 2: Overview of AlignOPSD . Cross-rollout alignment establishes comparable supervisory contexts, while correspondence changes define decision spans for hierarchical credit assignment. The weights enter policy optimization, with auxiliary components used during training.
ALFWorld
Search-QA
WebShop
Method
Pick
Heat
Look
Clean
Cool
Pick2
Avg
NQ
Triv
Pop
Hotp
2Wk
MuS
Bam
Avg
Score
Acc
Qwen2.5-3B-Instruct
Vanilla
44.4
15.4
11.1
6.2
28.6
12.5
21.9
24.6
48.1
31.0
26.3
25.3
7.2
59.7
31.7
6.7
0.8
Skill-Prompt ∗
51.7
0.0
66.7
48.4
4.3
10.0
28.9
23.7
46.2
30.6
24.4
22.1
7.5
12.5
23.9
1.2
0.8
OPSD
48.8
0.0
41.7
16.7
15.8
16.7
28.1
0.1
0.1
0.1
0.0
0.0
0.0
0.0
0.0
11.3
3.1
GRPO
91.2
61.9
62.5
96.2
65.0
47.4
75.0
39.3
60.6
41.1
37.4
34.6
15.4
26.4
36.4
79.8
63.3
Table 1: Performance on ALFWorld, Search-QA, and WebShop. We report success rate (%) on ALFWorld, accuracy (%) on Search-QA, and Score/Acc (%) on WebShop. Skills are training-only unless marked with ∗ (validation with skills). Best and second-best are highlighted.
#
Variant
Score
Acc
–
AlignOPSD
87.9
78.9
1
w/o rectification + adaptive allocation
84.2
71.9
2
w/ rectification + token allocation
81.6
71.1
3
w/ rectification + turn allocation
86.3
77.3
4
w/ rectification + random allocation
78.9
69.5
Table 2: Ablation Study. Final Score and Acc for AlignOPSD and four ablation variants are reported on the left and success-rate curves over 150 training steps are shown on the right.
Figure 3: Sensitivity Analysis. Sensitivity to retrieval top- K , thinking-similarity threshold γH , and credit temperature T across the tested hyperparameter ranges. Curves smooth recorded validation checkpoints, and shading shows twice the local temporal variation rather than across-seed uncertainty.
Figure 4: Mechanistic diagnostics. (a) Cross-rollout correspondence and confidence-weighted mixing; (b) correspondence-guided span segmentation and local versus rectified gaps; (c) task-dependent span-length distributions; (d) turn-averaged teacher–student gaps before and after rectification (left) and the distribution of absolute gap changes across target turns (right).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Interaction and outcome supervision
x,o0,kx
Task, initial observation, and training-only privileged information.
i,j;k,l;r;n
Trajectory, turn, token, and on-policy batch indices.
u=(i,k),v=(j,l)
Target turn and source turn from another same-task rollout.
τi,Ki
Trajectory and its number of turns.
hi,k,hi,k+
Ordinary history and privileged history (hi,k,kx) .
Appendix
Table 3: Notation used in AlignOPSD .
Statistic
S–S
S–P
Different action types at the same turn (%)
58.90 [56.76, 61.13]
55.06 [52.95, 57.19]
Matched entity-action pairs
143
289
Entity-action coverage (%)
1.71
1.29
Matched actions at different turns (%)
79.72
81.31
Absolute displacement, median (turns)
4
4
Appendix
Table 4: WebShop trajectory comparisons. Same-turn rates use 7,867 S–S and 21,265 S–P positions with two parseable actions; brackets give 95% task-bootstrap intervals (5,000 draws). Entity-action coverage is the number of matched pairs divided by valid left-turn exposures across trajectory pairs.
Figure 5: Temporal misalignment in trajectories and training. (a–b) Dots show per-turn estimates; faint dots in (b) show individual matches. Lines are local smooths with 95% task-bootstrap bands. (c) The first 100 updates, with bands of one local standard deviation.
Pair set
Count
Cosine mean [95% CI]
Std.
Same state (student–Teacher diagonal)
139,255
0.7604 [0.7580, 0.7630]
0.0977
All same-task, cross-rollout pairs
11,153,391
0.7214 [0.7192, 0.7237]
0.1014
Cross-rollout pairs above 0.8
2,564,616
0.8430 [0.8426, 0.8434]
0.0321
Cross-rollout pairs used by training
110,793
0.8805 [0.8791, 0.8820]
0.0418
Appendix
Table 5: Training-time thinking correspondence on WebShop over updates 1–100. Means and standard deviations pool recorded pairs; brackets are 95% update-bootstrap intervals for the mean. The diagonal count denotes canonical state exposures.
Figure 6: Span-length distributions on shared axes. Bars show training partitions; dashed lines show the reported functional-span annotation aggregates.
Setting
Value
Frozen encoder
Qwen3-Embedding-0.6B
Similarity threshold γH
0.80
Maximum sources K
3
Sources per sibling rollout
at most 1
Aggregation temperature / mode
0.10 / probability mixture
Maximum interpolation αmax
0.80
Appendix
Table 6: Rectification settings used to construct decision-aligned supervision.
Setting
Value
Profile temperature
0.10
Boundary quantile
0.80
Threshold range
[0.01,0.10]
Minimum / maximum span length
2 / 8 turns
Credit temperature
0.5
Span–turn mixing coefficient
0.5
Appendix
Table 7: Credit assignment settings used in AlignOPSD .
Setting
ALFWorld
WebShop
Search-QA
NVIDIA A800 GPUs
8
2
4
Training updates
150
150
150
Tasks per training batch
16
16
128
Rollouts per task
8
8
8
Maximum prompt tokens
2,048
4,096
4,096
Maximum response tokens/turn
512
512
512
Appendix
Table 8: Benchmark-specific training configuration shared by AlignOPSD and matched baselines.
Figure 7: Teacher–student gap during training. Identity and rectified gaps, together with their pointwise correction Δδ=δrect−δid , for Qwen2.5–3B/7B AlignOPSD runs on ALFWorld, WebShop, and Search-QA.
Figure 8: Critic scores and episode rewards during training. Logged mean critic scores and episode rewards for Qwen2.5–3B/7B AlignOPSD runs on ALFWorld, WebShop, and Search-QA.
Figure 9: Allocation diagnostics during training. Columns show ALFWorld, WebShop, and Search-QA. For each backbone, the first row reports turns per span and the second reports turn-weight standard deviation. Y-axis labels appear only in the leftmost column and x-axis labels only in the bottom row; each panel keeps its own y-axis ticks.
Figure 10: Cross-method performance traces. Rows correspond to Qwen2.5–3B and Qwen2.5–7B; columns correspond to ALFWorld, Search-QA, and WebShop. ALFWorld uses validation success, Search-QA uses logged episode success rate (EMA-7; faint lines show unsmoothed measurements), and WebShop uses validation normalized score (solid) and exact success (dashed). Each curve ends at its last logged update; no missing segment is extrapolated.