Organizations: The Chinese University of Hong Kong, Shenzhen · Shanghai AI Laboratory · University of Science and Technology of China · Independent · Institute of Computing Technology, CAS
Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF learns to predict environmental responses while distilling guidance from a feedback-conditioned self-teacher into the policy. Our analysis reveals a mutually reinforcing mechanism: environmental feedback modeling strengthens hindsight supervision and policy learning, while self-distillation enhances the model's ability to model environmental feedback. With Qwen3-8B, SELF outperforms SDPO and GRPO by 6.4 and 4.1 percentage points in τ-bench success rate, and by 10.71 and 3.57 percentage points in AppWorld task goal completion, respectively. These results show that SELF uses environmental feedback more effectively within agentic self-distillation, improving agent capabilities.
Figures & tables
Figure 1: Overview of SELF. The current policy interacts with the environment to collect transitions (hi,ai,fi) . Each transition is used for two complementary objectives: hindsight self-distillation, where a feedback-conditioned EMA teacher supervises the student policy, and environmental feedback modeling, where the model predicts the realized feedback from the interaction history and action. The two objectives are jointly optimized to obtain the next-round policy.
Method
Sampling
Supervision
Feedback Source
Teacher
Env. Modeling
SFT
Off-policy
Dense
Teacher
External
No
GRPO
On-policy
Sparse
Verifier
None
No
Mem0
–
–
Environment
None
No
SDPO
On-policy
Dense
Environment
Self
No
SELF (Ours)
On-policy
Dense
Environment
Self
Yes
Table 1: Comparison of learning paradigms under the configurations considered in this work. Supervision describes the granularity of the learning signal. Teacher denotes the source of teacher supervision. Env. Modeling indicates whether an explicit objective predicts environmental responses from interaction histories and actions. Dashes indicate inapplicable attributes.
Method
Setting
τ -bench
AppWorld
SR ↑
Turns ↓
TGC ↑
SGC ↑
External Model Baselines
Qwen3-30B-A3B-Instruct
–
33.6
11.9
31.55
12.50
Qwen3-4B-Instruct
Base
–
27.2 ± 1.4
10.5 ± 0.3
16.67 ± 1.19
5.36 ± 0.62
OEL
Observation
30.6 ± 1.5
10.9 ± 0.4
19.05 ± 1.43
6.25 ± 0.74
Table 2: Performance on τ -bench and AppWorld. SR denotes the pass@4 success rate on τ -bench, and Turns denotes the average number of interaction turns. TGC and SGC denote task goal completion and scenario goal completion on AppWorld, respectively. Observation and Full denote baseline feedback settings, while Reward denotes reward-based training. SELF learns from agent-generated actions and environmental observations, and its rows are shaded light blue. Higher SR, TGC, and SGC and lower Turns are better. Within each backbone, excluding external model baselines, the best means are bold with pale gold shading, and the second-best means are underlined with pale pink shading. Missing results and inapplicable settings are denoted by –.
Figure 2: Performance comparison on τ -bench Retail and AppWorld. (a) Success rate (SR, pass@4) gains of SELF over selected comparators on τ -bench Retail. Open and filled circles denote SELF and comparator scores, respectively; annotations and line colors indicate gains in percentage points. The superscript ‡ denotes Qwen3-8B; unmarked methods use Qwen3-4B-Instruct unless otherwise specified. (b) Task goal completion (TGC) and scenario goal completion (SGC) on AppWorld. Colors distinguish methods, and marker shapes distinguish model backbones.
Figure 3: Ablation studies and general capability evaluation with Qwen3-4B-Instruct. (a) Student and privileged teacher success rates (SR, pass@4) and average interaction turns on τ -bench Retail across different values of λo . (b) Environmental feedback prediction accuracy on τ -bench Retail and AppWorld across different values of λa . (c) Performance on MMLU, MMLU-Pro, and IFEval; SELF † denotes the privileged teacher. Vertical dashed lines in (a) and (b) indicate default settings.
Setting
Student
Privileged teacher
SR ↑
Turns ↓
SR ↑
Turns ↓
Base
27.2 ± 1.4
10.5 ± 0.3
–
–
λo=0
29.4 ± 2.3
11.2 ± 0.5
30.1 ± 1.1
10.5 ± 0.2
λo=0.1
36.5 ± 0.9
10.8 ± 0.4
49.1 ± 1.8
12.1 ± 0.6
λo=0.2
38.7 ± 1.5
10.4 ± 0.4
46.7 ± 2.4
10.9 ± 0.4
λo=0.5
35.7 ± 2.1
12.5 ± 0.6
44.0 ± 1.3
10.7 ± 0.3
Table 3: Effect of the environmental feedback modeling weight λo on τ -bench Retail with Qwen3-4B-Instruct. SR is reported as pass@4 (%), and Turns denotes the average number of interaction turns. Bold indicates the best result in each column.
Setting
τ -bench Retail
AppWorld
SR ↑
Turns ↓
Acc. ↑
TGC ↑
SGC ↑
Acc. ↑
Base
27.2 ± 1.4
10.5 ± 0.3
32 ± 1.8
16.67 ± 1.19
5.36 ± 0.62
11 ± 1.3
λa=0
20.1 ± 2.1
11.2 ± 0.5
37 ± 2.6
5.52 ± 0.83
1.21 ± 0.34
9 ± 1.7
λa=0.1
31.5 ± 1.3
12.0 ± 0.6
55 ± 1.5
26.67 ± 2.07
8.62 ± 0.91
39 ± 2.4
λa=0.2
33.2 ± 1.8
10.7 ± 0.3
62 ± 2.2
29.91 ± 1.43
8.45 ± 1.16
47 ± 1.6
λa=0.5
36.5 ± 0.9
10.2 ± 0.2
68 ± 1.3
31.10 ± 1.86
9.73 ± 0.78
54 ± 2.1
Table 4: Effect of the self-distillation weight λa on τ -bench Retail and AppWorld with Qwen3-4B-Instruct. SR is reported as pass@4, and Turns denotes the average number of interaction turns. All metrics except Turns are reported as percentages.
Method
Setting
MMLU ↑
MMLU-Pro ↑
IFEval ↑
Base
–
59.60 ± 0.18
34.38 ± 0.27
82.62 ± 0.35
GRPO
Reward
59.61 ± 0.24
34.25 ± 0.36
82.19 ± 0.48
SDPO
Observation
57.96 ± 0.31
31.02 ± 0.42
81.17 ± 0.53
SELF
Observation
59.57 ± 0.22
34.49 ± 0.33
82.61 ± 0.41
SELF (teacher)
Observation
59.71 ± 0.20
34.92 ± 0.29
82.58 ± 0.38
Table 5: General capability evaluation on MMLU, MMLU-Pro, and IFEval using Qwen3-4B-Instruct. Results are mean ± standard deviation. Light and darker blue shading identify the SELF student and teacher, respectively. Higher is better; bold entries indicate the highest mean per column.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Qwen3-4B-Instruct
Qwen3-8B
Optimizer
AdamW
Learning rate
3×10−4
Training batch size
16
LoRA rank r
16
96
LoRA alpha
32
128
LoRA dropout
0.05
0.05
Appendix
Table 6: Default SELF training configuration. Backbone-specific LoRA settings are listed separately; other settings are shared.
Condition
Retail
AppWorld
SR ↑
Turns ↓
TGC ↑
SGC ↑
Base
27.1
10.7
17.00
5.56
GRPO
33.0
13.4
28.67
10.50
SELF (cross-rollout)
36.6
11.7
31.21
8.67
SELF (same-rollout)
38.6
10.3
32.16
12.44
Appendix
Table 7: Effect of coupling environmental feedback modeling and hindsight self-distillation on the same rollout with Qwen3-4B-Instruct. In the cross-rollout condition, feedback modeling uses intact transitions from independently collected rollouts, while hindsight self-distillation uses the original rollout. Results are from one run per condition. SR, TGC, and SGC are percentages, with higher values indicating better performance. Turns is interpreted alongside SR. Bold values indicate the best completion scores.
Method
GPU-hours
Wall time(h)
OEL
20.3
6.12
SDPO (Observation)
11.8
3.67
GRPO
41.2
12.46
SELF
16.4
4.93
Appendix
Table 8: Training costs for the main experiments in Table 1 using four NVIDIA A100 80 GB GPUs. Values are aggregated across the reported training runs and exclude final evaluation and additional ablations. Lower values indicate lower cost.
Figure 4: Illustrative prompting interfaces used in SELF. The student prompt promptstu generates the next action ai from the interaction history hi ; the teacher prompt prompttea additionally conditions on the realized environmental feedback fi as privileged hindsight information; and the environmental-feedback modeling prompt promptenv predicts fi from hi and the current action ai .