cs.AIOct 8, 2026

Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation

Authors: Hangxi Guo, Fengyuan Liu, Yue Wang, Yuhua Qi, Haoyi Xiong, Fei Sun, Mengnan Du

Organizations: The Chinese University of Hong Kong, Shenzhen · Shanghai AI Laboratory · University of Science and Technology of China · Independent · Institute of Computing Technology, CAS

Abstract

Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF learns to predict environmental responses while distilling guidance from a feedback-conditioned self-teacher into the policy. Our analysis reveals a mutually reinforcing mechanism: environmental feedback modeling strengthens hindsight supervision and policy learning, while self-distillation enhances the model's ability to model environmental feedback. With Qwen3-8B, SELF outperforms SDPO and GRPO by 6.4 and 4.1 percentage points in ττ-bench success rate, and by 10.71 and 3.57 percentage points in AppWorld task goal completion, respectively. These results show that SELF uses environmental feedback more effectively within agentic self-distillation, improving agent capabilities.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation

    Jun 10, 2026Haoran Liu, Yuwei Zhang, Xiyao Li +2Credit Assignment in RLOn-Policy Distillation

  2. What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents

    May 19, 2026Xiaozhe Li, Tianyi Lyu, Yang Li +6Reinforcement LearningCredit Assignment in RL