Organizations: The Chinese University of Hong Kong, Shenzhen · Shanghai AI Laboratory · University of Science and Technology of China · Independent · Institute of Computing Technology, CAS
Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF learns to predict environmental responses while distilling guidance from a feedback-conditioned self-teacher into the policy. Our analysis reveals a mutually reinforcing mechanism: environmental feedback modeling strengthens hindsight supervision and policy learning, while self-distillation enhances the model's ability to model environmental feedback. With Qwen3-8B, SELF outperforms SDPO and GRPO by 6.4 and 4.1 percentage points in τ-bench success rate, and by 10.71 and 3.57 percentage points in AppWorld task goal completion, respectively. These results show that SELF uses environmental feedback more effectively within agentic self-distillation, improving agent capabilities.
Figures & tables
Figure 1: Overview of SELF. The current policy interacts with the environment to collect transitions (hi,ai,fi) . Each transition is used for two complementary objectives: hindsight self-distillation, where a feedback-conditioned EMA teacher supervises the student policy, and environmental feedback modeling, where the model predicts the realized feedback from the interaction history and action. The two objectives are jointly optimized to obtain the next-round policy.
Method
Sampling
Supervision
Feedback Source
Teacher
Env. Modeling
SFT
Off-policy
Dense
Teacher
External
No
GRPO
On-policy
Sparse
Verifier
None
No
Mem0
–
–
Environment
None
No
SDPO
On-policy
Dense
Environment
Self
No
SELF (Ours)
On-policy
Dense
Environment
Self
Yes
Table 1: Comparison of learning paradigms under the configurations considered in this work. Supervision describes the granularity of the learning signal. Teacher denotes the source of teacher supervision. Env. Modeling indicates whether an explicit objective predicts environmental responses from interaction histories and actions. Dashes indicate inapplicable attributes.
Method
Setting
τ -bench
AppWorld
SR ↑
Turns ↓
TGC ↑
SGC ↑
External Model Baselines
Qwen3-30B-A3B-Instruct
–
33.6
11.9
31.55
12.50
Qwen3-4B-Instruct
Base
–
27.2 ± 1.4
10.5 ± 0.3
16.67 ± 1.19
5.36 ± 0.62
OEL
Observation
30.6 ± 1.5
10.9 ± 0.4
19.05 ± 1.43
6.25 ± 0.74
Table 2: Performance on τ -bench and AppWorld. SR denotes the pass@4 success rate on τ -bench, and Turns denotes the average number of interaction turns. TGC and SGC denote task goal completion and scenario goal completion on AppWorld, respectively. Observation and Full denote baseline feedback settings, while Reward denotes reward-based training. SELF learns from agent-generated actions and environmental observations, and its rows are shaded light blue. Higher SR, TGC, and SGC and lower Turns are better. Within each backbone, excluding external model baselines, the best means are bold with pale gold shading, and the second-best means are underlined with pale pink shading. Missing results and inapplicable settings are denoted by –.
Figure 2: Performance comparison on τ -bench Retail and AppWorld. (a) Success rate (SR, pass@4) gains of SELF over selected comparators on τ -bench Retail. Open and filled circles denote SELF and comparator scores, respectively; annotations and line colors indicate gains in percentage points. The superscript ‡ denotes Qwen3-8B; unmarked methods use Qwen3-4B-Instruct unless otherwise specified. (b) Task goal completion (TGC) and scenario goal completion (SGC) on AppWorld. Colors distinguish methods, and marker shapes distinguish model backbones.
Figure 3: Ablation studies and general capability evaluation with Qwen3-4B-Instruct. (a) Student and privileged teacher success rates (SR, pass@4) and average interaction turns on τ -bench Retail across different values of λo . (b) Environmental feedback prediction accuracy on τ -bench Retail and AppWorld across different values of λa . (c) Performance on MMLU, MMLU-Pro, and IFEval; SELF † denotes the privileged teacher. Vertical dashed lines in (a) and (b) indicate default settings.
Setting
Student
Privileged teacher
SR ↑
Turns ↓
SR ↑
Turns ↓
Base
27.2 ± 1.4
10.5 ± 0.3
–
–
λo=0
29.4 ± 2.3
11.2 ± 0.5
30.1 ± 1.1
10.5 ± 0.2
λo=0.1
36.5 ± 0.9
10.8 ± 0.4
49.1 ± 1.8
12.1 ± 0.6
λo=0.2
38.7 ± 1.5
10.4 ± 0.4
46.7 ± 2.4
10.9 ± 0.4
λo=0.5
35.7 ± 2.1
12.5 ± 0.6
44.0 ± 1.3
10.7 ± 0.3
Table 3: Effect of the environmental feedback modeling weight λo on τ -bench Retail with Qwen3-4B-Instruct. SR is reported as pass@4 (%), and Turns denotes the average number of interaction turns. Bold indicates the best result in each column.
Setting
τ -bench Retail
AppWorld
SR ↑
Turns ↓
Acc. ↑
TGC ↑
SGC ↑
Acc. ↑
Base
27.2 ± 1.4
10.5 ± 0.3
32 ± 1.8
16.67 ± 1.19
5.36 ± 0.62
11 ± 1.3
λa=0
20.1 ± 2.1
11.2 ± 0.5
37 ± 2.6
5.52 ± 0.83
1.21 ± 0.34
9 ± 1.7
λa=0.1
31.5 ± 1.3
12.0 ± 0.6
55 ± 1.5
26.67 ± 2.07
8.62 ± 0.91
39 ± 2.4
λa=0.2
33.2 ± 1.8
10.7 ± 0.3
62 ± 2.2
29.91 ± 1.43
8.45 ± 1.16
47 ± 1.6
λa=0.5
36.5 ± 0.9
10.2 ± 0.2
68 ± 1.3
31.10 ± 1.86
9.73 ± 0.78
54 ± 2.1
Table 4: Effect of the self-distillation weight λa on τ -bench Retail and AppWorld with Qwen3-4B-Instruct. SR is reported as pass@4, and Turns denotes the average number of interaction turns. All metrics except Turns are reported as percentages.
Method
Setting
MMLU ↑
MMLU-Pro ↑
IFEval ↑
Base
–
59.60 ± 0.18
34.38 ± 0.27
82.62 ± 0.35
GRPO
Reward
59.61 ± 0.24
34.25 ± 0.36
82.19 ± 0.48
SDPO
Observation
57.96 ± 0.31
31.02 ± 0.42
81.17 ± 0.53
SELF
Observation
59.57 ± 0.22
34.49 ± 0.33
82.61 ± 0.41
SELF (teacher)
Observation
59.71 ± 0.20
34.92 ± 0.29
82.58 ± 0.38
Table 5: General capability evaluation on MMLU, MMLU-Pro, and IFEval using Qwen3-4B-Instruct. Results are mean ± standard deviation. Light and darker blue shading identify the SELF student and teacher, respectively. Higher is better; bold entries indicate the highest mean per column.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Qwen3-4B-Instruct
Qwen3-8B
Optimizer
AdamW
Learning rate
3×10−4
Training batch size
16
LoRA rank r
16
96
LoRA alpha
32
128
LoRA dropout
0.05
0.05
Appendix
Table 6: Default SELF training configuration. Backbone-specific LoRA settings are listed separately; other settings are shared.
Condition
Retail
AppWorld
SR ↑
Turns ↓
TGC ↑
SGC ↑
Base
27.1
10.7
17.00
5.56
GRPO
33.0
13.4
28.67
10.50
SELF (cross-rollout)
36.6
11.7
31.21
8.67
SELF (same-rollout)
38.6
10.3
32.16
12.44
Appendix
Table 7: Effect of coupling environmental feedback modeling and hindsight self-distillation on the same rollout with Qwen3-4B-Instruct. In the cross-rollout condition, feedback modeling uses intact transitions from independently collected rollouts, while hindsight self-distillation uses the original rollout. Results are from one run per condition. SR, TGC, and SGC are percentages, with higher values indicating better performance. Turns is interpreted alongside SR. Bold values indicate the best completion scores.
Method
GPU-hours
Wall time(h)
OEL
20.3
6.12
SDPO (Observation)
11.8
3.67
GRPO
41.2
12.46
SELF
16.4
4.93
Appendix
Table 8: Training costs for the main experiments in Table 1 using four NVIDIA A100 80 GB GPUs. Values are aggregated across the reported training runs and exclude final evaluation and additional ablations. Lower values indicate lower cost.
Figure 4: Illustrative prompting interfaces used in SELF. The student prompt promptstu generates the next action ai from the interaction history hi ; the teacher prompt prompttea additionally conditions on the realized environmental feedback fi as privileged hindsight information; and the environmental-feedback modeling prompt promptenv predicts fi from hi and the current action ai .
Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds, but not which intermediate actions caused the outcome or how they should be corrected. Recent methods alleviate this issue by generating rewards or textual hints from turn-level action-output signals, or by using feedback-conditioned self-distillation. However, generating feedback at every turn is inefficient when many intermediate turns are already successful or neutral, and applying feedback at a fixed or misaligned turn often fails to supervise the actions that contributed to the failure. To bridge this gap, we propose HINT-SD, a targeted self-distillation framework that uses full-trajectory hindsight to select failure-relevant actions and applies feedback-conditioned distillation only on targeted action spans. Experiments on BFCL v3 and AppWorld show that our method improves over the dense per-turn feedback baseline by up to 18.80 percent while achieving 2.26× lower time per training step, suggesting that selecting where to distill is a key factor for both effective and efficient long-horizon agent training.
Reinforcement learning typically improves multi-turn agent capabilities through the terminal outcome of the trajectories, which makes it difficult to determine credit assignments for each intermediate turns. Recent on-policy self-distillation methods offer a promising alternative by converting privileged feedback into dense token-level supervision through a self-teacher. Our study is motivated by the unexpected performance degradation observed when naively extending this paradigm to multi-turn settings, which we attribute to a lack of alignment between privileged feedback, such as successful trajectories or terminal outcomes, and the student's current decision context. We introduce HERO, a hindsight-enhanced self-distillation framework that uses next environment observations as locally aligned feedback. After each rollout, HERO reflects on the completed interaction to convert each observation into a compact turn-level diagnosis, that captures actionable feedback about the original action such as its necessity, validity or failure cause. On TauBench and WebShop, HERO improves task success and reduces unnecessary turns over environment-feedback-only self-distillation and GRPO. It is especially effective under limited training turn budgets, where successful rollouts are rare and GRPO provides weak reward-contrast signals.
Haoran Liu, Yuwei Zhang, Xiyao Li +2
University of California, San Diego · Independent Researcher · University of California, Berkeley
Reinforcement learning can train LLM agents from sparse task rewards, but long-horizon credit assignment remains challenging: a single success-or-failure signal must be distributed across many actions. Existing methods rely on trajectory-level rewards or proxy signals, without fully leveraging per-step environmental feedback. Multi-turn agent settings are underexplored, where feedback can include error messages, page changes, observations, or reference trajectories. We systematically study five feedback sources and two insertion granularities and introduce SERL, a selective environment-reweighted learning framework. SERL uses the task reward to determine update direction, while environment feedback adjusts placement and magnitude, focusing on critical actions. On ALFWorld and WebShop, SERL achieves 90.0% and 80.1% success, outperforming strong RL and distillation baselines. Analysis shows that grounded, action-relevant feedback at meaningful points consistently outperforms indiscriminate use of longer or richer context.
Xiaozhe Li, Tianyi Lyu, Yang Li +6
1Tongji University · 2Shanghai AI Laboratory · 5Independent +2