Organizations: University of Chinese Academy of Sciences · Institute of Automation, Chinese Academy of Sciences · Peking University · Mininglamp Technology · Key Laboratory of Computing Power Network and Information Security, Ministry of Education; Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences) · Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science
Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's sampled responses under privileged training-time context. However, our diagnostics show that positive average agreement between outcome and hindsight feedback coexists with substantial local disagreement, raising the question of how to allocate influence between them at each decision. We introduce UniOPSD (Unified On-Policy Self-Distillation), which unifies these feedback sources through adaptive local credit arbitration. UniOPSD constructs comparable credit estimates from environmental returns and successful-peer hindsight at shared interaction anchors. Historical agreement determines the global mixing level, while current signal availability and relative precision adjust each source's influence at individual decisions. The episode-level outcome contribution is retained, and bounded token modulation refines the fused step credit for policy optimization. With Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct, UniOPSD achieves ALFWorld success rates of 82.8% and 83.6%, WebShop success rates of 75.0% and 82.0%, and Search-QA aggregate accuracies of 45.3% and 49.8%, respectively. On 3B WebShop, UniOPSD improves over SDAR by 7.0 percentage points. Our code is available at https://github.com/Zenghuang-Fu/Uniopsd
Figures & tables
Figure 1: UniOPSD arbitrates between outcome and hindsight step credit at matched anchors. Historical agreement sets a global mixing level, and local evidence adjusts each step’s weight. The episode term is retained explicitly; bounded token modulation bk,t=1−λm+λmwk,t refines the arbitrated step term before PPO. Privileged peer context is restricted to training.
ALFWorld
Search-QA
WebShop
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg
NQ
Triv
Pop
Hotp
2Wk
MuS
Bam
Avg
Score
Acc
Qwen2.5-3B-Instruct
Vanilla
44.4
11.1
6.2
15.4
28.6
12.5
21.9
24.6
48.1
31.0
26.3
25.3
7.2
59.7
31.7
6.7
0.8
Skill-Prompt ∗
51.7
66.7
48.4
0.0
4.3
10.0
28.9
23.7
46.2
30.6
24.4
22.1
7.5
12.5
23.9
0.2
0.8
OPSD
48.8
41.7
16.7
0.0
15.8
16.7
28.1
0.1
0.1
0.1
0.0
0.0
0.0
0.0
0.0
11.3
3.1
GRPO
91.2
62.5
96.2
61.9
65.0
47.4
75.0
39.3
60.6
41.1
37.4
34.6
15.4
26.4
36.4
79.8
63.3
Table 1: Performance comparison of UniOPSD and baselines on ALFWorld, Search-QA, and WebShop. ALFWorld reports success rate, Search-QA reports accuracy, and WebShop reports score and accuracy (all in %). ∗ denotes validation with skills. Purple/bold and blue mark the highest and second-highest values, respectively, within each model and metric; ties share a color.
Qwen2.5-3B-Instruct
Qwen2.5-7B-Instruct
Variant
ALFWorld
Search-QA
WebShop
ALFWorld
Search-QA
WebShop
w/o step grouping (GRPO)
75.0
36.4
63.3
81.2
42.0
72.6
w/o teacher (GiGPO)
76.6
39.8
70.3
79.7
45.6
75.0
w/o correlation ( invvar )
78.9
41.3
71.9
78.1
46.3
76.6
w/ frozen ρ
80.5
42.8
73.4
82.0
47.5
78.9
UniOPSD
82.8
45.3
75.0
83.6
49.8
82.0
Table 2: Component ablations using overall benchmark results (%). ALFWorld and WebShop report success rates; Search-QA reports aggregate accuracy. The UniOPSD row reproduces Table 1 for reference; the frozen variant fixes ρ0=0.25 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
1
Freeze the behavior checkpoint πb and sample trajectory groups for the current tasks. Record trajectory returns, discounted step returns, anchor assignments, response masks, and student log probabilities.
2
Select a successful peer for each task group. Give its extracted action sequence to failed trajectories only. Score the same sampled tokens with πb under this prefix and compute Eq. ( 3 ).
3
Construct AkE and the weighted AkS using Appendix A.2 . Compute dk over every row of each anchor and the scale s over T . Form AT using Eq. ( 4 ).
4
Construct E , T , and B . Compute the precision proxies and their separate median normalizations, with source-specific fallback populations.
5
Read the stored correlation state. Compute ρm , cˉm , and ck , including the all-outcome endpoint. Form the fused step value Fk .
6
If the current B yields a valid Pearson correlation, update ρˉ and nprev for the next iteration. This update does not change the already formed ck .
7
Compute the scheduled token weights and Ak,t from Eqs. ( 9 )–( 10 ), retaining the response mask.
Appendix
Table 3: One UniOPSD iteration. All teacher scoring and advantage construction are performed without gradients. ρˉ and nprev denote the stored historical state.
Parameter
Symbol
Value
Global sensitivity
u
0.5
Correlation EMA decay
β
0.9
Correlation shrinkage strength
κ
2.33
Derived averaging horizon
H=(1−β)−1
10
Initial token modulation
λ0
0.5
Token modulation decay duration
Tdecay
100 updates
Appendix
Table 4: UniOPSD parameter settings for credit-signal analysis. The horizon H is derived from the EMA decay; the numerical margin η and stability constant ε serve different roles despite having the same value.
Setting
ALFWorld
WebShop
Search-QA
Training tasks per batch
16
16
128
Rollouts per task
8
8
8
Validation batch size
128
128
512
Maximum interaction turns
50
15
4
History length
2
2
4
Maximum prompt tokens
2048
4096
4096
Appendix
Table 5: Training settings for credit-signal analysis at both model scales within each environment. Prompt and response lengths are token limits per model call. The PPO minibatch size counts flattened training rows rather than complete episodes.
Method
Student tokens (M, cumulative)
Teacher-stage time (s/update)
Training time (s/update)
SDAR
432.61
59.22
427.63
UniOPSD
407.76
22.11
407.43
Reduction (%)
5.7
62.7
4.7
Appendix
Table 6: Runtime comparison with a local SDAR reproduction on ALFWorld-3B over 150 retained updates. Student tokens total the prompt and response tokens in the training batches. Training time excludes validation and checkpoint saving. Reductions use SDAR as the reference.
Environment
Model
Correlation
Disagreement (%)
Availability (%)
Mean ck
ALFWorld
3B
0.315
37.8
32.8
0.897
ALFWorld
7B
0.308
36.8
27.0
0.928
WebShop
3B
0.302
26.9
10.1
0.976
WebShop
7B
0.272
30.0
12.9
0.974
Search-QA
3B
0.371
26.8
11.7
0.950
Search-QA
7B
0.446
20.6
9.9
0.949
Appendix
Table 7: Step-signal statistics across model scales and environments over updates 1–150. Availability is the fraction of rows in Bm ; disagreement is conditional on that set. The final column is the all-row mean environment mixing weight and consequently includes rows without a teacher.
ALFWorld Successful Interaction
Task: put a clean lettuce in countertop.
Earlier action: clean lettuce 1 with sinkbasin 1
Observation: You clean the lettuce 1 using the sinkbasin 1.
Earlier action: go to countertop 1
Observation: You arrive at countertop 1. …
<think> Given the current situation, I have already cleaned the lettuce 1 using the sinkbasin 1. The next step should be to move the cleaned lettuce 1 to the countertop 1. The countertop 1 contains various items, but the lettuce 1 is not currently there. I need to move the lettuce 1 to the countertop 1 to complete the task. </think>
Appendix
Table 8: ALFWorld: cleaning an object before placing it at the target.
WebShop Successful Interaction
Task: Find me machine wash men’s t-shirts with long sleeve with color: swamp, and size: 3x-large big tall, and price lower than 70.00 dollars.
<think> Based on the current observation, the search results show men’s t-shirts with the color ’swamp’ and size ’3x-large big tall’. The price range is 10.52to40.50. The task is to find a t-shirt within this price range. The current action of clicking on ’swamp’ and ’3x-large big tall’ is the first step. The next logical action would be to click on ’buy now’ to purchase the t-shirt. I will now click on ’buy now’ to initiate the purchase process. </think>
<action> click[buy now] </action>
Appendix
Table 9: WebShop: selecting the requested color and size before purchase.
Search-QA Successful Interaction
Question: Vocelli Pizza and Noble Roman’s serve which kind of food?
<search> Vocelli Pizza and Noble Roman’s serve which kind of food </search>
<information> Doc 1 (Title: “Noble Roman’s”) Noble Roman’s is a pizza company based in Indianapolis, Indiana. … </information>
<search> Vocelli Pizza food type </search>
<information> Doc 1 (Title: “Vocelli Pizza”) Vocelli Pizza (formerly Pizza Outlet) is a pizzeria based in Pittsburgh, Pennsylvania. … </information>
<think> From the information provided, we know that Noble Roman’s is a pizza company. Similarly, Vocelli Pizza is also described as a pizzeria. Therefore, both of these establishments serve pizza. </think>
Appendix
Table 10: Search-QA: retrieving evidence about both entities before answering.
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We propose AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning. AgentOPSD aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space. This yields a principled reweighting scheme that converts sparse outcome supervision into turn-level credit signals and identifies pivotal turns through the marginal belief revision between consecutive states. The method is fully compatible with standard policy optimization and requires neither an additional critic nor extra rollouts. We evaluate AgentOPSD on ALFWorld, WebShop, and Search-QA using Qwen2.5 models at two scales (3B and 7B). AgentOPSD outperforms GRPO and strong self-distillation baselines, achieving 89.1% success on ALFWorld with Qwen2.5-7B. Ablation studies attribute the gains to turn-level aggregation and history-dependent recursive belief updates.
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10
Tsinghua University · Meituan · Zhejiang University
Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated. On-Policy Self-Distillation (OPSD) addresses this by re-scoring generated tokens under a privileged replay view to obtain dense, token-level supervision. However, we identify a confounding issue: the resulting support may reflect both the privileged information contained in the replay view and score shifts induced by the replay scaffold, making it difficult to attribute the support specifically to that information. This issue is especially pronounced when future environment observations serve as privileged information, since replaying them requires reconstructing an extended scaffold that itself perturbs token scores. To resolve this confounding, we propose Observation-Calibrated Self-Distillation (OCSD), which contrasts two structurally matched replay views, Full and Observation-Ablated, differing only in whether the actual future observation is present, to derive an observation residual that discounts score changes shared by the replay scaffold. OCSD then applies this residual to modulate token-level GRPO updates at high-uncertainty steps, while preserving the trajectory-level update direction. Experiments on ALFWorld, WebShop, and Search-QA across three Qwen3 model scales show that OCSD consistently outperforms strong baselines. Diagnostic analyses further confirm that the calibrated residual aligns better with local environment feedback. Our code is publicly available at https://github.com/yiy1x/OCSD.
Yi Yang, Cong Qin, Xiaodan Liu +8
Meituan LongCat Interaction · Nanjing University · Peking University +3
Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hinges on a few local decisions. Existing online policy distillation (OPD) provides denser token-level supervision, but typically treats heterogeneous agent trajectories as monolithic strings rather than causal interaction units. We present StepOPSD, a post-rollout preference self-distillation framework that takes the agent step as the unit of credit redistribution. StepOPSD decomposes trajectories into action-centered step segments, rescoring them under hindsight-enriched teacher contexts and converting token-level log-probability gaps into sign-preserving advantage shaping with a normalized per-step credit budget before the GRPO update. Across ALFWorld and Search-QA with Qwen3-1.7B and Qwen2.5-3B-Instruct, StepOPSD attains best or second-best results on subsets most sensitive to local causal errors, including first-place performance on ALFWorld Heat (79.1%), PickTwo (95.0%), Search-QA TriviaQA (61.6%), and tied-best performance on HotpotQA (40.4%). The results further reveal a consistent two-knob law: smaller α_clip acts as a broadly stabilizing local trust region, whereas the optimal global mixing strength λ_mix remains task-dependent. These findings suggest that step-aware distillation is most useful when trajectory-level rewards are weakly aligned with the local action that determines downstream success.