Organizations: University of Chinese Academy of Sciences · Institute of Automation, Chinese Academy of Sciences · Peking University · Mininglamp Technology · Key Laboratory of Computing Power Network and Information Security, Ministry of Education; Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences) · Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science
Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's sampled responses under privileged training-time context. However, our diagnostics show that positive average agreement between outcome and hindsight feedback coexists with substantial local disagreement, raising the question of how to allocate influence between them at each decision. We introduce UniOPSD (Unified On-Policy Self-Distillation), which unifies these feedback sources through adaptive local credit arbitration. UniOPSD constructs comparable credit estimates from environmental returns and successful-peer hindsight at shared interaction anchors. Historical agreement determines the global mixing level, while current signal availability and relative precision adjust each source's influence at individual decisions. The episode-level outcome contribution is retained, and bounded token modulation refines the fused step credit for policy optimization. With Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct, UniOPSD achieves ALFWorld success rates of 82.8% and 83.6%, WebShop success rates of 75.0% and 82.0%, and Search-QA aggregate accuracies of 45.3% and 49.8%, respectively. On 3B WebShop, UniOPSD improves over SDAR by 7.0 percentage points. Our code is available at https://github.com/Zenghuang-Fu/Uniopsd
Figures & tables
Figure 1: UniOPSD arbitrates between outcome and hindsight step credit at matched anchors. Historical agreement sets a global mixing level, and local evidence adjusts each step’s weight. The episode term is retained explicitly; bounded token modulation bk,t=1−λm+λmwk,t refines the arbitrated step term before PPO. Privileged peer context is restricted to training.
ALFWorld
Search-QA
WebShop
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg
NQ
Triv
Pop
Hotp
2Wk
MuS
Bam
Avg
Score
Acc
Qwen2.5-3B-Instruct
Vanilla
44.4
11.1
6.2
15.4
28.6
12.5
21.9
24.6
48.1
31.0
26.3
25.3
7.2
59.7
31.7
6.7
0.8
Skill-Prompt ∗
51.7
66.7
48.4
0.0
4.3
10.0
28.9
23.7
46.2
30.6
24.4
22.1
7.5
12.5
23.9
0.2
0.8
OPSD
48.8
41.7
16.7
0.0
15.8
16.7
28.1
0.1
0.1
0.1
0.0
0.0
0.0
0.0
0.0
11.3
3.1
GRPO
91.2
62.5
96.2
61.9
65.0
47.4
75.0
39.3
60.6
41.1
37.4
34.6
15.4
26.4
36.4
79.8
63.3
Table 1: Performance comparison of UniOPSD and baselines on ALFWorld, Search-QA, and WebShop. ALFWorld reports success rate, Search-QA reports accuracy, and WebShop reports score and accuracy (all in %). ∗ denotes validation with skills. Purple/bold and blue mark the highest and second-highest values, respectively, within each model and metric; ties share a color.
Qwen2.5-3B-Instruct
Qwen2.5-7B-Instruct
Variant
ALFWorld
Search-QA
WebShop
ALFWorld
Search-QA
WebShop
w/o step grouping (GRPO)
75.0
36.4
63.3
81.2
42.0
72.6
w/o teacher (GiGPO)
76.6
39.8
70.3
79.7
45.6
75.0
w/o correlation ( invvar )
78.9
41.3
71.9
78.1
46.3
76.6
w/ frozen ρ
80.5
42.8
73.4
82.0
47.5
78.9
UniOPSD
82.8
45.3
75.0
83.6
49.8
82.0
Table 2: Component ablations using overall benchmark results (%). ALFWorld and WebShop report success rates; Search-QA reports aggregate accuracy. The UniOPSD row reproduces Table 1 for reference; the frozen variant fixes ρ0=0.25 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
1
Freeze the behavior checkpoint πb and sample trajectory groups for the current tasks. Record trajectory returns, discounted step returns, anchor assignments, response masks, and student log probabilities.
2
Select a successful peer for each task group. Give its extracted action sequence to failed trajectories only. Score the same sampled tokens with πb under this prefix and compute Eq. ( 3 ).
3
Construct AkE and the weighted AkS using Appendix A.2 . Compute dk over every row of each anchor and the scale s over T . Form AT using Eq. ( 4 ).
4
Construct E , T , and B . Compute the precision proxies and their separate median normalizations, with source-specific fallback populations.
5
Read the stored correlation state. Compute ρm , cˉm , and ck , including the all-outcome endpoint. Form the fused step value Fk .
6
If the current B yields a valid Pearson correlation, update ρˉ and nprev for the next iteration. This update does not change the already formed ck .
7
Compute the scheduled token weights and Ak,t from Eqs. ( 9 )–( 10 ), retaining the response mask.
Appendix
Table 3: One UniOPSD iteration. All teacher scoring and advantage construction are performed without gradients. ρˉ and nprev denote the stored historical state.
Parameter
Symbol
Value
Global sensitivity
u
0.5
Correlation EMA decay
β
0.9
Correlation shrinkage strength
κ
2.33
Derived averaging horizon
H=(1−β)−1
10
Initial token modulation
λ0
0.5
Token modulation decay duration
Tdecay
100 updates
Appendix
Table 4: UniOPSD parameter settings for credit-signal analysis. The horizon H is derived from the EMA decay; the numerical margin η and stability constant ε serve different roles despite having the same value.
Setting
ALFWorld
WebShop
Search-QA
Training tasks per batch
16
16
128
Rollouts per task
8
8
8
Validation batch size
128
128
512
Maximum interaction turns
50
15
4
History length
2
2
4
Maximum prompt tokens
2048
4096
4096
Appendix
Table 5: Training settings for credit-signal analysis at both model scales within each environment. Prompt and response lengths are token limits per model call. The PPO minibatch size counts flattened training rows rather than complete episodes.
Method
Student tokens (M, cumulative)
Teacher-stage time (s/update)
Training time (s/update)
SDAR
432.61
59.22
427.63
UniOPSD
407.76
22.11
407.43
Reduction (%)
5.7
62.7
4.7
Appendix
Table 6: Runtime comparison with a local SDAR reproduction on ALFWorld-3B over 150 retained updates. Student tokens total the prompt and response tokens in the training batches. Training time excludes validation and checkpoint saving. Reductions use SDAR as the reference.
Environment
Model
Correlation
Disagreement (%)
Availability (%)
Mean ck
ALFWorld
3B
0.315
37.8
32.8
0.897
ALFWorld
7B
0.308
36.8
27.0
0.928
WebShop
3B
0.302
26.9
10.1
0.976
WebShop
7B
0.272
30.0
12.9
0.974
Search-QA
3B
0.371
26.8
11.7
0.950
Search-QA
7B
0.446
20.6
9.9
0.949
Appendix
Table 7: Step-signal statistics across model scales and environments over updates 1–150. Availability is the fraction of rows in Bm ; disagreement is conditional on that set. The final column is the all-row mean environment mixing weight and consequently includes rows without a teacher.
ALFWorld Successful Interaction
Task: put a clean lettuce in countertop.
Earlier action: clean lettuce 1 with sinkbasin 1
Observation: You clean the lettuce 1 using the sinkbasin 1.
Earlier action: go to countertop 1
Observation: You arrive at countertop 1. …
<think> Given the current situation, I have already cleaned the lettuce 1 using the sinkbasin 1. The next step should be to move the cleaned lettuce 1 to the countertop 1. The countertop 1 contains various items, but the lettuce 1 is not currently there. I need to move the lettuce 1 to the countertop 1 to complete the task. </think>
Appendix
Table 8: ALFWorld: cleaning an object before placing it at the target.
WebShop Successful Interaction
Task: Find me machine wash men’s t-shirts with long sleeve with color: swamp, and size: 3x-large big tall, and price lower than 70.00 dollars.
<think> Based on the current observation, the search results show men’s t-shirts with the color ’swamp’ and size ’3x-large big tall’. The price range is 10.52to40.50. The task is to find a t-shirt within this price range. The current action of clicking on ’swamp’ and ’3x-large big tall’ is the first step. The next logical action would be to click on ’buy now’ to purchase the t-shirt. I will now click on ’buy now’ to initiate the purchase process. </think>
<action> click[buy now] </action>
Appendix
Table 9: WebShop: selecting the requested color and size before purchase.
Search-QA Successful Interaction
Question: Vocelli Pizza and Noble Roman’s serve which kind of food?
<search> Vocelli Pizza and Noble Roman’s serve which kind of food </search>
<information> Doc 1 (Title: “Noble Roman’s”) Noble Roman’s is a pizza company based in Indianapolis, Indiana. … </information>
<search> Vocelli Pizza food type </search>
<information> Doc 1 (Title: “Vocelli Pizza”) Vocelli Pizza (formerly Pizza Outlet) is a pizzeria based in Pittsburgh, Pennsylvania. … </information>
<think> From the information provided, we know that Noble Roman’s is a pizza company. Similarly, Vocelli Pizza is also described as a pizzeria. Therefore, both of these establishments serve pizza. </think>
Appendix
Table 10: Search-QA: retrieving evidence about both entities before answering.