Recent advancements in Large Language Model (LLM) agents have demonstrated strong capabilities in executing complex tasks through tool use. However, long-horizon multi-step tool planning is challenging, because the exploration space suffers from a combinatorial explosion. In this scenario, even when a correct tool-use path is found, it is usually considered an immediate reward for current training, which would not provide any reusable information for subsequent training. In this paper, we argue that historically successful trajectories contain reusable tool-transition patterns, which can be leveraged throughout the whole training process. Inspired by ant colony optimization where historically successful paths can be reflected by the pheromone, we propose Pheromone-Guided Policy Optimization (PhGPO), which learns a trajectory-based transition pattern (i.e., pheromone) from historical trajectories and then uses the learned pheromone to guide policy optimization. This learned pheromone provides explicit and reusable guidance that steers policy optimization toward historically successful tool transitions, thereby improving long-horizon tool planning. Comprehensive experimental results demonstrate the effectiveness of our proposed PhGPO.
Figures & tables
Figure 1: Two-stage pilot experiment with an explicit transition prior for long-horizon tool planning. We distill cross-trajectory transition memory from verified successful tool-use trajectories in an easier stage, then reuse it as a prior for GRPO optimization. Panels (a,b) show higher average return and success rate under matched interaction budgets, (c) fewer steps to the first successful tool-use trajectory, and (d) larger relative gains with longer trajectories.
Figure 2: Overview of PhGPO. PhGPO converts verified successful tool-use trajectories into a pheromone-based explicit transition prior over tool-transition and tool-to-invocation edges. The prior is updated through deposition and evaporation, and incorporated into rollout generation to guide the policy toward historically successful transitions and argument invocations. Training proceeds from supervised next-tool warm-up to progressive pheromone-guided reinforcement learning and finally full pheromone-guided policy optimization.
Method
Toolathlon
TRAJECT-Bench
TOUCAN
Qwen2.5-7B
Llama3.1-8B
Qwen2.5-7B
Llama3.1-8B
Qwen2.5-7B
Llama3.1-8B
Match R.
TSR
Match R.
TSR
Match R.
TSR
Match R.
TSR
Match R.
TSR
Match R.
TSR
Standard Prompting & Planning Strategies
ReAct
5.31
2.92
5.15
2.48
15.32
9.20
9.16
7.01
11.64
8.92
7.03
5.88
Plan-and-Solve
6.42
4.61
4.27
3.96
12.44
11.35
9.05
5.54
12.71
10.15
7.42
5.12
Beyond ReAct
3.85
3.10
5.23
4.18
11.33
10.74
15.61
14.20
12.92
10.84
7.35
6.54
Table 1: Main Results on Long-Horizon Tool-Use Benchmarks. We report Match Ratio (Match R., %) and Task Success Rate (TSR, %) on Toolathlon, TRAJECT-Bench, and TOUCAN across two backbone models. Higher values ( ↑ ) indicate better performance.
Figure 3: Emergence of the reference chain in pheromone transitions. As training proceeds, the transition matrix becomes increasingly concentrated on reference edges (red), while unrelated transitions fade due to evaporation, and pheromone values on the reference chain increase steadily.
Variant
Variant role
Match R. ↑
Drop M. ↓
TSR ↑
Drop TSR ↓
w/o Pheromone ( β=0 throughout)
removes explicit transition prior
21.23
4.02
16.72
3.62
w/o Mixed Curriculum
removes progressive oracle guidance
19.34
5.91
14.86
5.48
Static Prior
freezes pheromone updates
23.93
1.32
18.71
1.63
w/o Evaporation ( ρ=0 )
removes decay of stale statistics
21.72
3.53
16.28
4.06
w/o Task-dependent Pheromones
removes task-matched edge memory
22.45
2.80
17.35
2.99
PhGPO (Full)
all components enabled
25.25
–
20.34
–
Table 2: Ablation studies on Toolathlon. We report Match Ratio (Match R., %) and Task Success Rate (TSR, %) on the test set. Drop columns are computed relative to PhGPO (Full), where larger drops indicate stronger contribution of the removed component.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Tool source
L1 Tools
L2 Tools
#Instances
Avg Steps
Toolathlon
MCP apps
618
1,250
108 tasks
13.48
TRAJECT-Bench
executable APIs
381
715
5,870 queries
6.36
TOUCAN (Multi-step)
MCP servers
850
15,294
1,646,546 traj.
3.16
Appendix
Table 3: Dataset Statistics. We summarize the scale of the three benchmarks under our Tool-Transition Graph action representation. L1 Tools denotes the number of base tools, and L2 Tools denotes the number of argument-invocation patterns. Avg Steps reports the average number of tool calls per verifiably successful reference tool-use sequence in the retained episodes. For TOUCAN, we only consider Multi-step instances.
Figure 4: Ablation study of pheromone influence parameter β on Toolathlon benchmark. (a) Learning curves show that dynamic β annealing achieves the highest Match Ratio (25.25%), significantly outperforming fixed strategies. β=0 (no pheromone) exhibits slow, noisy convergence, while β=5 (over-guidance) becomes trapped in local optima. (b) Next-tool accuracy mirrors learning trends, with dynamic β reaching 27.16% through effective balance of exploration and exploitation. (c) Exploration diversity exhibits natural fluctuations due to finite-sample estimation from rollouts, but reveals clear trends: β=0 maintains excessively high diversity (unfocused exploration), β=5 shows extremely low diversity (trapped exploitation), while dynamic β demonstrates an optimal trajectory—starting with high exploration (0.75) and smoothly transitioning into the optimal range (0.55-0.65) that balances verified pattern reuse with adaptive exploration.
Setting
Match R.
TAcc
Final Diversity
β=0 (No pheromone)
21.23
19.86
0.76
β=1 (Fixed)
23.93
22.05
0.53
β=5 (Over-guidance)
17.80
16.50
0.25
Dynamic β (Ours)
25.25
27.16
0.58
Appendix
Table 4: Ablation on β for Toolathlon (Qwen2.5-7B). We report Match Ratio (Match R., %), Next-tool Accuracy (TAcc, %), and final Exploration Diversity.
Figure 5: Training Dynamics of PhGPO. (a) Average Return : All backbone models show a consistent upward trend in average return, indicating stable policy improvement during training. (b) Pheromone Graph Growth : The number of discovered edges grows rapidly in early epochs and stabilizes later, reflecting fast discovery of feasible tool transitions followed by refinement on a stable set of edges.
Figure 6: Hyperparameter sensitivity analysis on Toolathlon. We evaluate 8 hyperparameters for pheromone guidance and policy optimization. Red stars indicate our selected values; shaded regions show ±1 standard deviation across 3 seeds. Row 1: (a) βmax ; (b) ρ ; (c) α ; (d) w . Row 2: (e) M ; (f) η ; (g) ptf ; (h) γ .
Category
Parameter
Value
Description
Pheromone
βmax
0.7
Maximum pheromone influence weight in Eq. ( 7 )
ρ
0.01
Evaporation rate in Eq. ( 1 )
α
1.0
Deposition rate in Eq. ( 1 )
w
0.5
Task-dependent pheromone weight in Eq. ( 5 )
Optimization
M
5
GRPO group size, measured as rollouts per instance
η
10−7
Learning rate for RL fine-tuning
Appendix
Table 5: Key hyperparameters in PhGPO and selected values.
Optimizer variant
Qwen2.5-7B
Llama3.1-8B
Match R. ↑
TAcc ↑
Match R. ↑
TAcc ↑
PhGPO + GRPO
25.25
27.16
22.83
23.09
w/o pheromone guidance ( β=0 )
21.23
19.86
18.75
16.45
PhGPO + PPO
23.27
25.03
18.21
24.05
w/o pheromone guidance ( β=0 )
20.78
23.31
15.72
21.22
PhGPO + RLOO
20.21
24.47
18.51
22.69
Appendix
Table 6: PhGPO with different policy optimization backbones on Toolathlon. For each backbone optimizer, we report results with pheromone-guided sampling (Eq. ( 7 )) and without pheromone guidance by setting β=0 in sampling. Higher is better ( ↑ ).
Figure 7: Pheromone evolution on a longer tool-use trajectory, visualized on a 10-tool subset. Top: fused tool-transition pheromone τTool(⋅∣x) restricted to ten tools at checkpoint, with reference-chain edges highlighted in red. Bottom: pheromone values on the highlighted reference-chain transitions at the same checkpoint. The visualization shows that pheromone concentrates on reference-chain transitions while off-chain transitions fade under evaporation, indicating that the explicit transition prior remains selective on longer tool-use trajectories.
Figure 8: Time efficiency analysis on Toolathlon. We track the earliest training step at which a high-quality trajectory ( qτ(ξ)≥0.6 ) is first generated for each episode. PhGPO reaches broad coverage earlier than the ablation without pheromone, and the relative improvement remains broadly consistent across trajectory-length bins.
Method
Train step (s)
Train overhead
Peak GPU (GB)
Infer/step (ms)
Infer overhead
GRPO
11.86
baseline
17.3
58.5
baseline
+ graph maintenance
11.87
+0.01s
17.3
58.8
+0.3ms
+ graph + pheromone sampling
11.88
+0.02s
17.4
59.1
+0.6ms
Full PhGPO
11.89
+0.03s
17.4
59.3
+0.8ms
Appendix
Table 7: Runtime and memory overhead on Toolathlon. We report end-to-end overhead using Qwen2.5-7B.
Setting
β schedule
prand
Match R. ↑
TAcc ↑
Final Diversity
No pheromone
0
0.05
21.23
19.86
0.76
Full model w/o random exploration
dynamic β
0
24.63
26.34
0.52
Full model
dynamic β
0.05
25.25
27.16
0.58
Full model, larger random exploration
dynamic β
0.10
24.38
26.47
0.66
Appendix
Table 8: Random exploration and pheromone guidance on Toolathlon. We compare the effects of removing pheromone guidance and changing the random exploration probability prand using Qwen2.5-7B.
LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizons. However, existing benchmarks rarely evaluate planning under retrieval-limited tool visibility. To address this gap, we introduce PlanBench-XL, an interactive benchmark of 327 retail tasks over 1,665 tools that tests whether agents can iteratively retrieve usable tools, invoke them to uncover intermediate evidence for subsequent calls toward the final goal. PlanBench-XL further features an optional blocking mechanism that simulates real-world unpredictability through missing, failing, or distracting tool functions, forcing agents to detect disrupted paths and adapt at runtime. Experiments on ten leading LLMs show that massive-tool planning remains challenging: while GPT-5.4 achieves 51.90% accuracy in block-free settings, it collapses to 11.36% under the most severe blocking condition. Further analysis shows that agents are especially vulnerable when failures lack explicit error signals or when recovery requires longer alternative tool-use paths. These results establish PlanBench-XL as a testbed for diagnosing agentic planning failures and highlight the need for robust adaptive planning in long-horizon tasks with large, imperfect tool environments.
While Large Language Models (LLMs) have demonstrated strong capabilities as autonomous agents across a wide range of tasks, their performance often degrades in multi-turn long-horizon agentic tasks. Existing methods have made progress through fine-grained credit assignment to alleviate long-horizon sparse rewards and hierarchical reinforcement learning to decompose tasks and reduce long-term dependency. However, these methods still do not directly address long-context interference, in which continuously growing histories weaken the agent's ability to track the global task state and impair subsequent reasoning and decision-making. Inspired by the way humans handle complex tasks through subgoal decomposition and completed progress summarization, we propose Hierarchical Planning and Information Folding (HIPIF) for long-horizon LLM agent learning. HIPIF trains the agent end-to-end to organize long-horizon execution around explicit subgoals while folding completed subgoal histories to reduce long-context interference. Furthermore, to stabilize subgoal-based planning and execution, HIPIF combines hierarchical reflection and subgoal-oriented process rewards to guide subgoal generation, transition, and execution, without relying on costly auxiliary models or task-specific expert trajectories. Extensive experiments on three publicly available agentic benchmarks demonstrate the validity of our method.
Juncheng Diao, Zhicong Lu, Peiguang Li +6
1Meituan · University of Chinese Academy of Sciences
Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Existing approaches directly construct tool-level graphs from these trajectories, but the resulting graphs remain tied to specific tools and are hard to generalize across tool sets. To tackle this challenge, we find that despite differences in the tools involved, analogous tasks often share a common function-level workflow structure, which serves as a potentially more transferable abstraction for tool planning. Based on this insight, we propose ToolLIFT, a framework that lifts tool-specific trajectories into a function-level workflow graph (FWG) for generalizable tool planning. Specifically, we first propose a trajectory-lifting mechanism that encodes workflow structures in the FWG and shares collaboration experience across tools. Then, building on the global structure of the FWG, we introduce decoupled workflow planning and tool selection to align individual tool choices with the overall workflow. Lastly, to ensure reliable tool dataflow, we adopt Reinforcement Learning (RL) and propose source-gated and skill-specific rewards to maintain source-traceable information flow across tool calls. Experiments on two in-distribution (ID) and three out-of-distribution (OOD) benchmarks show that ToolLIFT consistently outperforms state-of-the-art baselines, demonstrating strong generalization to unseen tool sets.
Xiuhui You, Jiayi Luo, Zichao Shen +2
School of Computer Science and Engineering · Beihang University