We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost.
Figures & tables
Figure 1: A flow diagram of how STRAT gets the required labels without human intervention.
Family
Difficulty
Definition
Placed
Easy
No rule produces the goal tile, so it is already on the grid. The goal is reach or hold.
Go-hold
Medium
The goal tile is crafted. The goal is reach or hold, and every earlier step is only walk or pick up.
Beside
Hard
The goal tile is crafted. The goal is “put A beside B ,” and every earlier step is only walk or pick up.
Align
Hardest
The goal tile is crafted, and some earlier step must line two objects up to produce a tile.
Table 1: Ruleset structure families. A task takes the first matching row. Placed is decided by whether any rule produces a goal tile; the other three are read from the backward chain, whose last phase is the goal.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
small-1m
medium-1m / high-1m
Environment
Environment
XLand-MiniGrid-R1-9x9
Observation
symbolic 5×5 egocentric (tile, colour) + heading
Actions
6
Episode length
243
Reward
1−0.9t/Tmax on success, else 0
Appendix
Table 2: Full hyperparameter settings. Values spanning both columns are shared by all benchmarks.
Large language models (LLMs) are increasingly used as interactive agents, but optimizing them for long-horizon decision making remains difficult because current methods are largely purely reactive, which weakens both exploration and credit assignment over extended trajectories. In this work, we present Strategic Trajectory Abstraction (StraTA), a simple framework that introduces an explicit trajectory-level strategy into agentic reinforcement learning (RL). StraTA samples a compact strategy from the initial task state, conditions subsequent actions on that strategy, and trains strategy generation and action execution jointly with a hierarchical GRPO-style rollout design, further enhanced by diverse strategy rollout and critical self-judgment. Experiments on ALFWorld, WebShop, and SciWorld show that StraTA consistently improves both sample efficiency and final performance over strong baselines. StraTA reaches success rates of 93.1% on ALFWorld and 84.2% on WebShop. On SciWorld, StraTA attains a 63.5% overall score, outperforming frontier closed-source models.
Xiangyuan Xue, Yifan Zhou, Zidong Wang +5
1The Chinese University of Hong Kong · 2Shanghai Artificial Intelligence Laboratory · University of Georgia +2
Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training. Outcome rewards provide reliable supervision for short-horizon reasoning, but become sparse and high-variance as trajectories grow to tens or hundreds of tool calls. They can also be misleading: a failed rollout may contain many useful actions that move the agent closer to the goal, yet outcome-only training assigns them the same negative advantage as the eventual mistake. We propose TRACE (Turn-level Reward Assignment via Credit Estimation), a dense credit-assignment method for agentic reinforcement learning. TRACE represents rollouts as state transitions at tool-call boundaries, obtains gold-answer log-probabilities from a frozen reference model, transforms them into log-ratio state values, and derives per-action rewards as Temporal-Difference changes in those values. This requires no additional critic or process-label training, and its one-step log-ratio TD component telescopes across redundant tool calls. On long-horizon complex search, TRACE substantially improves base-model tool-use ability using pure RL, without a cold-start supervised fine-tuning stage, an agentic mid-training stage, or training on live-web data. On the closed-web BrowseComp-Plus benchmark, it raises Qwen3-4B from 7.2 to 35.6 and Qwen3-30B-A3B from 8.4 to 42.6. The learned search behavior also transfers to open-web benchmarks, and the learning curves show earlier improvement and faster convergence during RL training.
Leitian Tao, Baolin Peng, Wenlin Yao +5
University of Wisconsin–Madison · Microsoft Research
Motivated by the challenge presented by non-Markovian objectives in reinforcement learning (RL), we present a novel framework to track and represent the progress of autonomous agents through complex, multi-stage tasks. Given a specification in finite linear temporal logic (LTL), the framework establishes a 'tracking vector' which updates at each time step in a trajectory rollout. The values of the vector represent the status of the specification as the trajectory develops, assigning true, false, or 'open' labels (where 'open' is used for indeterminate cases). Applied to an LTL formula tree, the tracking vector can be used to encode detailed information about how a task is executed over a trajectory, providing a potential tool for new performance metrics, diverse exploration, and reward shaping. In this paper, we formally present the framework and algorithm, collectively named Live LTL Progress Tracking, give a simple working example, and demonstrate avenues for its integration into RL models. Future work will apply the framework to problems such as task-space exploration and diverse solution-finding in RL.
Noel Brindise, Cedric Langbort, Melkior Ornik
Coordinated Science Laboratory, University of Illinois Urbana-Champaign, Urbana, USA