State Trace Rationale As Auxiliary Task in Reinforcement Learning
Organizations: University of the Witwatersrand Johannesburg, South Africa · New York University New York, USA
Abstract
We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost.
Figures & tables
| Family | Difficulty | Definition |
|---|---|---|
| Placed | Easy | No rule produces the goal tile, so it is already on the grid. The goal is reach or hold. |
| Go-hold | Medium | The goal tile is crafted. The goal is reach or hold, and every earlier step is only walk or pick up. |
| Beside | Hard | The goal tile is crafted. The goal is “put beside ,” and every earlier step is only walk or pick up. |
| Align | Hardest | The goal tile is crafted, and some earlier step must line two objects up to produce a tile. |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| small-1m | medium-1m / high-1m | |
| Environment | ||
| Environment | XLand-MiniGrid-R1-9x9 | |
| Observation | symbolic egocentric (tile, colour) + heading | |
| Actions | 6 | |
| Episode length | 243 | |
| Reward | on success, else 0 | |