Hindsight Experience Replay (HER) enhances sample efficiency in goal-conditioned reinforcement learning (RL) by relabelling failed trajectories with goals that were actually achieved. However, HER assumes access to a goal relabelling function and a predicate function that determines whether a goal has been satisfied. These assumptions break down in instruction-following tasks, where goals are expressed in natural language and differ from the state space. We formalize this as the Hindsight Reinforcement Learning problem, which shows the need to jointly learn these functions alongside the RL policy. To address it, we propose ETHER (Emergent Textual Hindsight Experience Replay), an agent that leverages Emergent Communication. ETHER uses a referential game (RG) to train a speaker and a listener to develop a grounded, artificial language describing environment states. It partially aligns this emergent language with instruction language using co-occurrence patterns between task instructions and RL observations. Experiments on BabyAI's PickupDist task show that ETHER's learned RG speaker and listener can function as the goal relabelling and predicate functions of HER, improving sample efficiency despite imperfect language alignment. Our work bridges Emergent Communication and goal-conditioned RL, opening the door to wider applications of HER.
Figures & tables
Figure 1 : (a): Schema of a discriminative object-centric/2-player/ L -signal/ N=0 -round/ K=64 -distractor visual referential game [ 22 ] using a Straight-Through Gumbel-Softmax (STGS) communication channel [ 27 ] , with a modified STGS loss adapted to LazImpa [ 65 ] , termed STGS-LazImpa (see Appendix F.1 ). Stimuli undergo data augmentation from [ 23 ] to enforce object-centricism [ 17 , 22 ] , differing between speaker and listener. (b, c): Histograms of semantic counts from goal-conditioned trajectories by BabyAI’s expert agent and a Random agent, showing that goal semantics align with the most frequently observed features.
Figure 2 : ETHER uses three agents, an RL agent and a speaker and a listener from a RG. Failed trajectories, where the RL agent did not follow instructions g , are relabelled by the RG speaker with an alternative goal g^=mRG(sT) . The RG listener , within the predicate function fRG , adjusts the rewards throughout the trajectory to reflect the extent to which each state match the alternative goal g^ . The relabelled trajectory is then added to the replay buffer, from which samples can be used to update all agents with losses JHRL and JRG .
Agent
Success %
Any-Colour
Any-Shape
R2D2
58.01±4.07
–
–
Pickup+Descr HER
61.72±3.91
100.00±0.0
100.0±0.0
Pickup HER
74.41±6.83
6.24±1.48
6.24±1.48
ETHER
67.68±4.56
24.76±5.30
15.95±22.56
Table 1 : Performances (mean ± std.err.) of different agents after 1 M observations at the instruction-following task, and language alignment’s accuracies in terms of the Any-Shape and Any-Colour metrics.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
R2D2
Number of actors
32
Actor update interval
1 env. step
Sequence unroll length
20
Sequence length overlap
10
Sequence burn-in length
10
N-steps return
3
Appendix
Table 2 : Hyper-parameter values relevant to R2D2 in the different architectures presented. Missing parameters follow Ape-X [ 33 ] .
Agent
Mean
R2D2 (no Burn-In)
13.02 ± 1.26
HIGhER+ (no Burn-In)
16.02 ± 1.79
HIGhER++ (n=1) (no Burn-In)
14.97 ± 1.19
HIGhER++ (n=2) (no Burn-In)
15.89 ± 0.60
HIGhER++ (n=4) (no Burn-In)
13.93 ± 2.29
Appendix
Table 3 : Success ratios (mean and standard deviation) for agents without the burn-in feature of R2D2 after 200k steps in a modified version of the BabyAI PickUpDist-v0 task. 3 random seeds for each agent.
Figure 3 : Left: Trajectories for the blue colour goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 4 : Left: Trajectories for colorless objects from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 5 : Left: Trajectories for the green color goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 6 : Left: Trajectories for the grey color goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 7 : Left: Trajectories for the purple color goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 8 : Left: Trajectories for the red color goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.
Figure 9 : Left: Trajectories for the yellow color goal from BabyAI’s built-in expert agent which always reaches the goal. Right: Random agent trajectories. In both cases the semantics of the goal are among the most observed semantic features for any given trajectory. This effect is less pronounced in the random agent.