cs.AISep 29, 2026

State Trace Rationale As Auxiliary Task in Reinforcement Learning

Authors: Muhammad U. Nasir, Alex Vogt, Steven D. James, Julian Togelius

Organizations: University of the Witwatersrand Johannesburg, South Africa · New York University New York, USA

Abstract

We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction

    May 7, 2026Xiangyuan Xue, Yifan Zhou, Zidong Wang +5Agentic Reinforcement LearningAgentic Learning

  2. TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

    Jul 15, 2026Leitian Tao, Baolin Peng, Wenlin Yao +5Credit AssignmentAgentic Reinforcement Learning

  3. Live LTL Progress Tracking: Towards Task-Based Exploration

    Apr 18, 2026Noel Brindise, Cedric Langbort, Melkior OrnikLinear Temporal LogicsTrajectory Generation