LLM agents often admit multiple high-quality solutions to the same task, differing in reasoning structure, tool-use pattern, or interaction trajectory. Yet existing notions of diversity in LLM post-training are mostly implicit, arising from general stochasticity and regularization mechanisms rather than explicitly targeting task-relevant behavioral variation. While such implicit diversity can be useful, it does not directly specify which forms of behavioral variation should be encouraged for a given task. In this work, we study explicit trajectory diversity in RL-based post-training for LLMs. Our key idea is to define diversity through user-specified, task-specific trajectory descriptors, which map each sampled trajectory to an interpretable behavioral representation, and then measure diversity as a set-level functional over the resulting descriptor matrix. Building on this formulation, we introduce Trajectory-guided Joint Policy Optimization(TJPO), a single-policy framework that optimizes explicit diversity over sampled trajectory groups, avoiding the need for population-based policy training, and instantiate it within group-based policy optimization through trajectory-level learning signals. This design makes the diversity objective both interpretable and controllable. Experiments on Sokoban and ALFWorld show that TJPO improves task-specific trajectory diversity while maintaining competitive task performance. Descriptor and trajectory analyses show that the learned variation follows the specified behavioral dimensions and includes distinct successful strategies. Extra experiment results suggest that explicitly shaping trajectory diversity can help LLM agents satisfy user requirements and remain effective when task conditions change.
Figures & tables
Figure 1: Motivation and schematic overview of TJPO. Without an explicit diversity objective, successful rollouts may concentrate on similar behaviors (top). TJPO augments the group-based advantage with descriptor-based trajectory-level diversity credits to encourage task-relevant behavioral variation (bottom).
Table 2
Figure 2: Successful navigation strategies for “put a spraybottle in cabinet”. TJPO uses the location descriptor with pairwise diversity and λ=0.3 . Both methods succeed in all eight rollouts. Nodes denote visited locations; edge labels and widths encode navigation transition counts across the eight rollouts.
Table 4
Method
Standard Success ↑
User Constraints Success ↑
Disruption Success ↑
GiGPO
0.9917
0.7708
0.7792
TJPO
0.9750
0.8542
0.8542
Table 5: Practical benefits of behavioral diversity on ALFWorld.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
ALFWorld
Sokoban
Base model
Qwen2.5-7B-Instruct
Qwen3-VL-4B-Instruct
Input modality
Text
Vision-language
Environment
AlfredTWEnv
Sokoban
Environment observation
Text
RGB image
Max environment steps
50
15
Training data size
16
32
Appendix
Table 6: Environment-specific experimental setup.
Hyperparameter
Value
Training framework
verl
RL backbone
GiGPO
Group size G
8
Max response length
512
Learning rate
1×10−6
KL loss type
low-variance KL
Appendix
Table 7: Shared training and evaluation hyperparameters.
Environment
Descriptor
What it captures
Vector type
ALFWorld
action_type
High-level action
normalized histogram ( R15 )
ALFWorld
location
Navigation destination usage
normalized histogram ( R∣L(x)∣ )
Sokoban
direction_transition
Consecutive direction pairs
normalized histogram ( R16 )
Sokoban
spatial_coverage
Visited grid-cell distribution
normalized 2D occupancy ( R7×7 )
Appendix
Table 8: Summary of trajectory descriptors used in our experiments.
Figure 3: Location-descriptor view of the main-paper route case study for the prompt “put a spraybottle in cabinet.” Each row is one successful rollout, each column is a location activated by at least one rollout, and color denotes the normalized frequency of navigation actions targeting that location. Rows are ordered by their projection onto the first principal component of the pooled GiGPO and TJPO descriptor vectors, and both panels use a shared color scale.
Figure 4: Complete successful environment-action trajectories for the same “put a spraybottle in cabinet” case. All eight sampled rollouts are shown for each method; no trajectories are omitted or selected. Rollouts are grouped by navigation strategy while retaining their original rollout identifiers. Object instance indices are omitted for readability, but every environment action is shown. Colors distinguish navigation destinations and task-completion operations.
Figure 5: Complete successful trajectories for “clean some ladle and put it in diningtable.” The top panel shows GiGPO and the bottom panel shows location-pairwise TJPO with λ=0.3 . All eight rollouts are retained for each method, with their original rollout identifiers preserved. Object instance indices and singleton-location indices are omitted for readability, while Countertop 1 and Countertop 2 are distinguished. Rows are grouped by navigation sequence; both panels use the same action-position scale, and no actions are omitted.
LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful branches from the same decision state. We study this problem as successful strategy coverage: how broadly a model realizes distinct successful strategies under a fixed rollout budget. We present Direct Diversity Optimization (DDO), an offline post-training method that combines Divergence-Tree Collection (DTC) with the Reference-Relative Target-Odds Objective (RTO). DTC constructs state-aligned branch sets rooted at shared decision states, and RTO trains the model to match reference-relative targets over successful alternatives. DDO achieves the strongest task success and successful strategy coverage among the compared post-training methods across BabyAI, BabaIsAI, and WebShop. It also achieves the highest recovery rate after local action replacement and higher task success and coverage than successful-only imitation and decoding-time diversification controls.
Junwon Ko, Dong-Jae Lee, Minchan Kwon +2
School of Electrical Engineering, KAIST Daejeon, Republic of Korea
RL post-training strategies are dataset-dependent and reveal a recurring empirical pattern: capacity parameters accumulate monotonically across stages, while regularization parameters predominantly oscillate in response to shifting training dynamics. This distinction highlights a potential flaw in fixed training schedules: by forcing all parameters along rigid paths, they fail to capture the dynamic exploration-exploitation tradeoffs that regularization must track. We uncover this through LLMZero, an agentic system that optimizes training trajectories via tree search by diagnosing pathologies at each checkpoint and proposing coordinated multi-parameter transitions. Across four diverse GRPO tasks, LLMZero discovers strategies that improve over the base model by 9% to 140% and over grid search by 6% to 15% (relative), consistently outperforming random search and a skill-based agent under a matched compute budget. The capacity--regularization asymmetry is consistent across all four tasks, offering a candidate design heuristic for multi-stage training.
Haoyang Fang, Wei Zhu, Boran Han +11
†LLMZero Project Core Team. · ∗Work done at Amazon.
Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed rewards often lead to trajectory neglect, in which agents lose focus on the task goal and interaction history at intermediate steps. Prior work has explored step-level supervision using Shannon-entropy-based uncertainty signals, which conflate inherent state complexity with agent confidence and therefore provide unreliable estimates of decision reliability. To address this issue, we propose normalized entropy, which measures confidence deviations relative to an agent's average behavior under a given state, thereby strengthening the association between low-quality actions and trajectory neglect. Building on this insight, we introduce Selective Trajectory-Aware Policy Optimization (STAPO), a hierarchical group-based RL framework. STAPO leverages normalized entropy to locate outlier steps associated with trajectory neglect and optimizes them via a joint mechanism of trajectory-aware reward and trajectory-independent penalty, enhancing trajectory awareness while preserving training stability. Extensive experiments on ALFWorld, WebShop, and Search-Augmented QA demonstrate that STAPO achieves state-of-the-art performance while substantially alleviating trajectory neglect, validating its effectiveness and robustness for agentic tasks.
Qiuyi Qi, Tian Liang, Mutian Bao +8
Zhejiang University · Ant Group · City University of Hong Kong