cs.CLJul 28, 2023

ETHER: Aligning Emergent Communication for Hindsight Experience Replay

Authors: Kevin Yandoka Denamganaï, Daniel Hernandez, Ozan Vardal, Sondess Missaoui, James Alfred Walker

Organizations: Department of Computer Science University of York York, UK

Abstract

Hindsight Experience Replay (HER) enhances sample efficiency in goal-conditioned reinforcement learning (RL) by relabelling failed trajectories with goals that were actually achieved. However, HER assumes access to a goal relabelling function and a predicate function that determines whether a goal has been satisfied. These assumptions break down in instruction-following tasks, where goals are expressed in natural language and differ from the state space. We formalize this as the Hindsight Reinforcement Learning problem, which shows the need to jointly learn these functions alongside the RL policy. To address it, we propose ETHER (Emergent Textual Hindsight Experience Replay), an agent that leverages Emergent Communication. ETHER uses a referential game (RG) to train a speaker and a listener to develop a grounded, artificial language describing environment states. It partially aligns this emergent language with instruction language using co-occurrence patterns between task instructions and RL observations. Experiments on BabyAI's PickupDist task show that ETHER's learned RG speaker and listener can function as the goal relabelling and predicate functions of HER, improving sample efficiency despite imperfect language alignment. Our work bridges Emergent Communication and goal-conditioned RL, opening the door to wider applications of HER.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Select-to-Act: Hierarchical Reinforcement Learning via Adaptive Language Guidance

    Jun 21, 2026Hanping Zhang, Adam Koziak, Yuhong GuoHierarchical Reinforcement LearningEnglish

  2. Learning More from Less: Reinforcement Learning from Hindsight

    Jul 10, 2026Iris Xu, Sunshine Jiang, John Marangola +8Reinforcement Learning Post-TrainingHindsight