Hindsight Experience Replay (HER) enhances sample efficiency in goal-conditioned reinforcement learning (RL) by relabelling failed trajectories with goals that were actually achieved. However, HER assumes access to a goal relabelling function and a predicate function that determines whether a goal has been satisfied. These assumptions break down in instruction-following tasks, where goals are expressed in natural language and differ from the state space. We formalize this as the Hindsight Reinforcement Learning problem, which shows the need to jointly learn these functions alongside the RL policy. To address it, we propose ETHER (Emergent Textual Hindsight Experience Replay), an agent that leverages Emergent Communication. ETHER uses a referential game (RG) to train a speaker and a listener to develop a grounded, artificial language describing environment states. It partially aligns this emergent language with instruction language using co-occurrence patterns between task instructions and RL observations. Experiments on BabyAI's PickupDist task show that ETHER's learned RG speaker and listener can function as the goal relabelling and predicate functions of HER, improving sample efficiency despite imperfect language alignment. Our work bridges Emergent Communication and goal-conditioned RL, opening the door to wider applications of HER.
Figures & tables
Figure 1 : (a): Schema of a discriminative object-centric/2-player/ L -signal/ N=0 -round/ K=64 -distractor visual referential game [ 22 ] using a Straight-Through Gumbel-Softmax (STGS) communication channel [ 27 ] , with a modified STGS loss adapted to LazImpa [ 65 ] , termed STGS-LazImpa (see Appendix F.1 ). Stimuli undergo data augmentation from [ 23 ] to enforce object-centricism [ 17 , 22 ] , differing between speaker and listener. (b, c): Histograms of semantic counts from goal-conditioned trajectories by BabyAI’s expert agent and a Random agent, showing that goal semantics align with the most frequently observed features.
Figure 2 : ETHER uses three agents, an RL agent and a speaker and a listener from a RG. Failed trajectories, where the RL agent did not follow instructions g , are relabelled by the RG speaker with an alternative goal g^=mRG(sT) . The RG listener , within the predicate function fRG , adjusts the rewards throughout the trajectory to reflect the extent to which each state match the alternative goal g^ . The relabelled trajectory is then added to the replay buffer, from which samples can be used to update all agents with losses JHRL and JRG .
Agent
Success %
Any-Colour
Any-Shape
R2D2
58.01±4.07
–
–
Pickup+Descr HER
61.72±3.91
100.00±0.0
100.0±0.0
Pickup HER
74.41±6.83
6.24±1.48
6.24±1.48
ETHER
67.68±4.56
24.76±5.30
15.95±22.56
Table 1 : Performances (mean ± std.err.) of different agents after 1 M observations at the instruction-following task, and language alignment’s accuracies in terms of the Any-Shape and Any-Colour metrics.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
R2D2
Number of actors
32
Actor update interval
1 env. step
Sequence unroll length
20
Sequence length overlap
10
Sequence burn-in length
10
N-steps return
3
Appendix
Table 2 : Hyper-parameter values relevant to R2D2 in the different architectures presented. Missing parameters follow Ape-X [ 33 ] .
Agent
Mean
R2D2 (no Burn-In)
13.02 ± 1.26
HIGhER+ (no Burn-In)
16.02 ± 1.79
HIGhER++ (n=1) (no Burn-In)
14.97 ± 1.19
HIGhER++ (n=2) (no Burn-In)
15.89 ± 0.60
HIGhER++ (n=4) (no Burn-In)
13.93 ± 2.29
Appendix
Table 3 : Success ratios (mean and standard deviation) for agents without the burn-in feature of R2D2 after 200k steps in a modified version of the BabyAI PickUpDist-v0 task. 3 random seeds for each agent.
Figure 3 : Left: Trajectories for the blue colour goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 4 : Left: Trajectories for colorless objects from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 5 : Left: Trajectories for the green color goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 6 : Left: Trajectories for the grey color goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 7 : Left: Trajectories for the purple color goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 8 : Left: Trajectories for the red color goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 9 : Left: Trajectories for the yellow color goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Reinforcement Learning (RL) has been widely applied to sequential decision-making, yet it often suffers from poor sample efficiency due to costly interactions with the environment. A limited line of recent work has started exploring improving RL efficiency by leveraging external knowledge expressed in natural-language instructions. However, the few existing approaches typically treat the entire instruction as a single conditioning input, failing to account for the stage-dependent nature of language guidance, especially in complex environments. In this paper, we propose \emph{Hierarchical Reinforcement Learning with Language Instructions (HRLLI)}, a hierarchical RL framework that explicitly models natural-language instructions as dynamically selectable semantic guidance during decision-making. HRLLI decomposes instructions into a set of piecewise guidance elements, where each instruction piece may become relevant at different stages of interaction with the environment. A novel hierarchical RL policy structure is then formulated in a \emph{Select-to-Act} paradigm: a high-level semantic policy acts as a guidance selector that selects the most relevant instruction piece to the current state to guide the low-level agent's decision, while a low-level policy executes environment actions conditioned on the selected guidance. The two-level policies are learned simultaneously to maximize augmented expected returns from interactions with the environment. This design enables the agent to adaptively ground language instructions into stage-specific decisions during interaction. Experiments on the instruction-intensive RTFM benchmark show that HRLLI consistently outperforms strong instruction-conditioned RL baselines, demonstrating that explicitly modeling adaptive instruction selection significantly improves the effectiveness of RL.
Hanping Zhang, Adam Koziak, Yuhong Guo
School of Computer Science, Carleton University, Ottawa, Canada · Canada CIFAR AI Chair, Amii, Canada
Large language model agents operate in partially observable, long-horizon settings where obtaining supervision remains a major bottleneck. We address this by utilizing a source of supervision overlooked in existing post-training methods: unintended yet successful goals embedded within agent rollouts. Specifically, we introduce Hindsight Supervised Learning (HSL), where an auxiliary LLM reviews each completed trajectory and relabels it with all of the natural-language goals the agent actually achieved. HSL then pairs the trajectory with its relabeled goals and uses these pairs for additional fine-tuning. To mitigate suboptimality in the relabeled data, we propose two learning techniques for HSL, irrelevant-action masking and sample reweighting. Our experiments show that HSL is flexible and compatible with existing post-training pipelines. It improves both SFT and DPO, with larger gains on long-horizon tasks with more diverse goal spaces. Moreover, HSL is sample-efficient: on ALFWorld, it surpasses baselines trained on the full dataset while using only one quarter of the ground-truth demonstrations.
Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern. Manipulation tasks typically provide only sparse rewards, so a weak policy fails almost every rollout early in training and has little to learn from, even when those failures execute coherent behavior. Such a failure, however, is a success at a different task. We present Learning from Hindsight (LfH), which brings hindsight relabeling to RL post-training of VLAs by scoring failed rollouts against the tasks they actually achieved. A single vision-language model relabels both the instruction and the reward, proposing a hindsight instruction for a group of failed rollouts and scoring how well each satisfies it, and the policy trains on the relabeled and original rollouts jointly. Because VLAs generalize across language, relabeling in language lets the policy learn more from the same trajectories. On out-of-distribution LIBERO-PRO tasks, where standard RL improves only slowly, LfH achieves 5× improvement in sample efficiency, and outperforms a dense progress-reward baseline. The gains hold across VLA backbones and on a physical Franka robot.
Iris Xu, Sunshine Jiang, John Marangola +8
1Massachusetts Institute of Technology · 2MIT-IBM Computing Research Lab · 3Stanford University +1