Organizations: Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University · School of Artificial Intelligence, Beihang University · Tsinghua University
Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 % and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD
Figures & tables
Figure 1: Decision–Timestamp Mismatch. Functional decisions need not share timestamps across different sibling rollouts and can span multiple turns within individual trajectories.
Figure 2: Overview of AlignOPSD . Cross-rollout alignment establishes comparable supervisory contexts, while correspondence changes define decision spans for hierarchical credit assignment. The weights enter policy optimization, with auxiliary components used during training.
ALFWorld
Search-QA
WebShop
Method
Pick
Heat
Look
Clean
Cool
Pick2
Avg
NQ
Triv
Pop
Hotp
2Wk
MuS
Bam
Avg
Score
Acc
Qwen2.5-3B-Instruct
Vanilla
44.4
15.4
11.1
6.2
28.6
12.5
21.9
24.6
48.1
31.0
26.3
25.3
7.2
59.7
31.7
6.7
0.8
Skill-Prompt ∗
51.7
0.0
66.7
48.4
4.3
10.0
28.9
23.7
46.2
30.6
24.4
22.1
7.5
12.5
23.9
1.2
0.8
OPSD
48.8
0.0
41.7
16.7
15.8
16.7
28.1
0.1
0.1
0.1
0.0
0.0
0.0
0.0
0.0
11.3
3.1
GRPO
91.2
61.9
62.5
96.2
65.0
47.4
75.0
39.3
60.6
41.1
37.4
34.6
15.4
26.4
36.4
79.8
63.3
Table 1: Performance on ALFWorld, Search-QA, and WebShop. We report success rate (%) on ALFWorld, accuracy (%) on Search-QA, and Score/Acc (%) on WebShop. Skills are training-only unless marked with ∗ (validation with skills). Best and second-best are highlighted.
#
Variant
Score
Acc
–
AlignOPSD
87.9
78.9
1
w/o rectification + adaptive allocation
84.2
71.9
2
w/ rectification + token allocation
81.6
71.1
3
w/ rectification + turn allocation
86.3
77.3
4
w/ rectification + random allocation
78.9
69.5
Table 2: Ablation Study. Final Score and Acc for AlignOPSD and four ablation variants are reported on the left and success-rate curves over 150 training steps are shown on the right.
Figure 3: Sensitivity Analysis. Sensitivity to retrieval top- K , thinking-similarity threshold γH , and credit temperature T across the tested hyperparameter ranges. Curves smooth recorded validation checkpoints, and shading shows twice the local temporal variation rather than across-seed uncertainty.
Figure 4: Mechanistic diagnostics. (a) Cross-rollout correspondence and confidence-weighted mixing; (b) correspondence-guided span segmentation and local versus rectified gaps; (c) task-dependent span-length distributions; (d) turn-averaged teacher–student gaps before and after rectification (left) and the distribution of absolute gap changes across target turns (right).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Interaction and outcome supervision
x,o0,kx
Task, initial observation, and training-only privileged information.
i,j;k,l;r;n
Trajectory, turn, token, and on-policy batch indices.
u=(i,k),v=(j,l)
Target turn and source turn from another same-task rollout.
τi,Ki
Trajectory and its number of turns.
hi,k,hi,k+
Ordinary history and privileged history (hi,k,kx) .
Appendix
Table 3: Notation used in AlignOPSD .
Statistic
S–S
S–P
Different action types at the same turn (%)
58.90 [56.76, 61.13]
55.06 [52.95, 57.19]
Matched entity-action pairs
143
289
Entity-action coverage (%)
1.71
1.29
Matched actions at different turns (%)
79.72
81.31
Absolute displacement, median (turns)
4
4
Appendix
Table 4: WebShop trajectory comparisons. Same-turn rates use 7,867 S–S and 21,265 S–P positions with two parseable actions; brackets give 95% task-bootstrap intervals (5,000 draws). Entity-action coverage is the number of matched pairs divided by valid left-turn exposures across trajectory pairs.
Figure 5: Temporal misalignment in trajectories and training. (a–b) Dots show per-turn estimates; faint dots in (b) show individual matches. Lines are local smooths with 95% task-bootstrap bands. (c) The first 100 updates, with bands of one local standard deviation.
Pair set
Count
Cosine mean [95% CI]
Std.
Same state (student–Teacher diagonal)
139,255
0.7604 [0.7580, 0.7630]
0.0977
All same-task, cross-rollout pairs
11,153,391
0.7214 [0.7192, 0.7237]
0.1014
Cross-rollout pairs above 0.8
2,564,616
0.8430 [0.8426, 0.8434]
0.0321
Cross-rollout pairs used by training
110,793
0.8805 [0.8791, 0.8820]
0.0418
Appendix
Table 5: Training-time thinking correspondence on WebShop over updates 1–100. Means and standard deviations pool recorded pairs; brackets are 95% update-bootstrap intervals for the mean. The diagonal count denotes canonical state exposures.
Figure 6: Span-length distributions on shared axes. Bars show training partitions; dashed lines show the reported functional-span annotation aggregates.
Setting
Value
Frozen encoder
Qwen3-Embedding-0.6B
Similarity threshold γH
0.80
Maximum sources K
3
Sources per sibling rollout
at most 1
Aggregation temperature / mode
0.10 / probability mixture
Maximum interpolation αmax
0.80
Appendix
Table 6: Rectification settings used to construct decision-aligned supervision.
Setting
Value
Profile temperature
0.10
Boundary quantile
0.80
Threshold range
[0.01,0.10]
Minimum / maximum span length
2 / 8 turns
Credit temperature
0.5
Span–turn mixing coefficient
0.5
Appendix
Table 7: Credit assignment settings used in AlignOPSD .
Setting
ALFWorld
WebShop
Search-QA
NVIDIA A800 GPUs
8
2
4
Training updates
150
150
150
Tasks per training batch
16
16
128
Rollouts per task
8
8
8
Maximum prompt tokens
2,048
4,096
4,096
Maximum response tokens/turn
512
512
512
Appendix
Table 8: Benchmark-specific training configuration shared by AlignOPSD and matched baselines.
Figure 7: Teacher–student gap during training. Identity and rectified gaps, together with their pointwise correction Δδ=δrect−δid , for Qwen2.5–3B/7B AlignOPSD runs on ALFWorld, WebShop, and Search-QA.
Figure 8: Critic scores and episode rewards during training. Logged mean critic scores and episode rewards for Qwen2.5–3B/7B AlignOPSD runs on ALFWorld, WebShop, and Search-QA.
Figure 9: Allocation diagnostics during training. Columns show ALFWorld, WebShop, and Search-QA. For each backbone, the first row reports turns per span and the second reports turn-weight standard deviation. Y-axis labels appear only in the leftmost column and x-axis labels only in the bottom row; each panel keeps its own y-axis ticks.
Figure 10: Cross-method performance traces. Rows correspond to Qwen2.5–3B and Qwen2.5–7B; columns correspond to ALFWorld, Search-QA, and WebShop. ALFWorld uses validation success, Search-QA uses logged episode success rate (EMA-7; faint lines show unsmoothed measurements), and WebShop uses validation normalized score (solid) and exact success (dashed). Each curve ends at its last logged update; no missing segment is extrapolated.
On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need not improve future behavior, while consequential guidance may be beyond the current student's reach or fail to persist without privileged input. We formulate long-horizon OPD as hierarchical supervision allocation and argue that productive guidance lies at the intersection of future utility and current learnability. Crucially, this intersection evolves as the student learns. Based on this principle, we propose LENS-OPD, a coarse-to-fine framework that organizes supervision through Locate, Validate, and Refine. Locate adapts trajectory exposure to the student's evolving competence and proposes a candidate decision for intervention. Validate tests whether teacher guidance at that decision improves the same student's subsequent behavior. Refine internalizes the beneficial guided behavior into the deployable policy and concentrates token-level supervision on decisive teacher-student conflicts within the validated turn. These stages are nested: each finer allocation is conditioned on the coarser decision, rather than being optimized as an independent importance score. Experiments across multiple long-horizon agent benchmarks and student-teacher configurations show that LENS-OPD consistently improves task performance over vanilla OPD and strong curriculum- and selection-based baselines. Our results suggest that effective long-horizon distillation requires teaching at the right depth, the right decision, and the right token.
Yuhao Sun, Binrui Wu, Zhuoer Xu +5
Ant Group · University of Science and Technology of China · Alibaba International Digital Commerce Group +2
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We propose AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning. AgentOPSD aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space. This yields a principled reweighting scheme that converts sparse outcome supervision into turn-level credit signals and identifies pivotal turns through the marginal belief revision between consecutive states. The method is fully compatible with standard policy optimization and requires neither an additional critic nor extra rollouts. We evaluate AgentOPSD on ALFWorld, WebShop, and Search-QA using Qwen2.5 models at two scales (3B and 7B). AgentOPSD outperforms GRPO and strong self-distillation baselines, achieving 89.1% success on ALFWorld with Qwen2.5-7B. Ablation studies attribute the gains to turn-level aggregation and history-dependent recursive belief updates.
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10
Tsinghua University · Meituan · Zhejiang University
Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hinges on a few local decisions. Existing online policy distillation (OPD) provides denser token-level supervision, but typically treats heterogeneous agent trajectories as monolithic strings rather than causal interaction units. We present StepOPSD, a post-rollout preference self-distillation framework that takes the agent step as the unit of credit redistribution. StepOPSD decomposes trajectories into action-centered step segments, rescoring them under hindsight-enriched teacher contexts and converting token-level log-probability gaps into sign-preserving advantage shaping with a normalized per-step credit budget before the GRPO update. Across ALFWorld and Search-QA with Qwen3-1.7B and Qwen2.5-3B-Instruct, StepOPSD attains best or second-best results on subsets most sensitive to local causal errors, including first-place performance on ALFWorld Heat (79.1%), PickTwo (95.0%), Search-QA TriviaQA (61.6%), and tied-best performance on HotpotQA (40.4%). The results further reveal a consistent two-knob law: smaller α_clip acts as a broadly stabilizing local trust region, whereas the optimal global mixing strength λ_mix remains task-dependent. These findings suggest that step-aware distillation is most useful when trajectory-level rewards are weakly aligned with the local action that determines downstream success.