LLM agents often admit multiple high-quality solutions to the same task, differing in reasoning structure, tool-use pattern, or interaction trajectory. Yet existing notions of diversity in LLM post-training are mostly implicit, arising from general stochasticity and regularization mechanisms rather than explicitly targeting task-relevant behavioral variation. While such implicit diversity can be useful, it does not directly specify which forms of behavioral variation should be encouraged for a given task. In this work, we study explicit trajectory diversity in RL-based post-training for LLMs. Our key idea is to define diversity through user-specified, task-specific trajectory descriptors, which map each sampled trajectory to an interpretable behavioral representation, and then measure diversity as a set-level functional over the resulting descriptor matrix. Building on this formulation, we introduce Trajectory-guided Joint Policy Optimization(TJPO), a single-policy framework that optimizes explicit diversity over sampled trajectory groups, avoiding the need for population-based policy training, and instantiate it within group-based policy optimization through trajectory-level learning signals. This design makes the diversity objective both interpretable and controllable. Experiments on Sokoban and ALFWorld show that TJPO improves task-specific trajectory diversity while maintaining competitive task performance. Descriptor and trajectory analyses show that the learned variation follows the specified behavioral dimensions and includes distinct successful strategies. Extra experiment results suggest that explicitly shaping trajectory diversity can help LLM agents satisfy user requirements and remain effective when task conditions change.
Figures & tables
Figure 1: Motivation and schematic overview of TJPO. Without an explicit diversity objective, successful rollouts may concentrate on similar behaviors (top). TJPO augments the group-based advantage with descriptor-based trajectory-level diversity credits to encourage task-relevant behavioral variation (bottom).
Table 2
Figure 2: Successful navigation strategies for “put a spraybottle in cabinet”. TJPO uses the location descriptor with pairwise diversity and λ=0.3 . Both methods succeed in all eight rollouts. Nodes denote visited locations; edge labels and widths encode navigation transition counts across the eight rollouts.
Table 4
Method
Standard Success ↑
User Constraints Success ↑
Disruption Success ↑
GiGPO
0.9917
0.7708
0.7792
TJPO
0.9750
0.8542
0.8542
Table 5: Practical benefits of behavioral diversity on ALFWorld.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
ALFWorld
Sokoban
Base model
Qwen2.5-7B-Instruct
Qwen3-VL-4B-Instruct
Input modality
Text
Vision-language
Environment
AlfredTWEnv
Sokoban
Environment observation
Text
RGB image
Max environment steps
50
15
Training data size
16
32
Appendix
Table 6: Environment-specific experimental setup.
Hyperparameter
Value
Training framework
verl
RL backbone
GiGPO
Group size G
8
Max response length
512
Learning rate
1×10−6
KL loss type
low-variance KL
Appendix
Table 7: Shared training and evaluation hyperparameters.
Environment
Descriptor
What it captures
Vector type
ALFWorld
action_type
High-level action
normalized histogram ( R15 )
ALFWorld
location
Navigation destination usage
normalized histogram ( R∣L(x)∣ )
Sokoban
direction_transition
Consecutive direction pairs
normalized histogram ( R16 )
Sokoban
spatial_coverage
Visited grid-cell distribution
normalized 2D occupancy ( R7×7 )
Appendix
Table 8: Summary of trajectory descriptors used in our experiments.
Figure 3: Location-descriptor view of the main-paper route case study for the prompt “put a spraybottle in cabinet.” Each row is one successful rollout, each column is a location activated by at least one rollout, and color denotes the normalized frequency of navigation actions targeting that location. Rows are ordered by their projection onto the first principal component of the pooled GiGPO and TJPO descriptor vectors, and both panels use a shared color scale.
Figure 4: Complete successful environment-action trajectories for the same “put a spraybottle in cabinet” case. All eight sampled rollouts are shown for each method; no trajectories are omitted or selected. Rollouts are grouped by navigation strategy while retaining their original rollout identifiers. Object instance indices are omitted for readability, but every environment action is shown. Colors distinguish navigation destinations and task-completion operations.
Figure 5: Complete successful trajectories for “clean some ladle and put it in diningtable.” The top panel shows GiGPO and the bottom panel shows location-pairwise TJPO with λ=0.3 . All eight rollouts are retained for each method, with their original rollout identifiers preserved. Object instance indices and singleton-location indices are omitted for readability, while Countertop 1 and Countertop 2 are distinguished. Rows are grouped by navigation sequence; both panels use the same action-position scale, and no actions are omitted.