Recent advancements in Large Language Model (LLM) agents have demonstrated strong capabilities in executing complex tasks through tool use. However, long-horizon multi-step tool planning is challenging, because the exploration space suffers from a combinatorial explosion. In this scenario, even when a correct tool-use path is found, it is usually considered an immediate reward for current training, which would not provide any reusable information for subsequent training. In this paper, we argue that historically successful trajectories contain reusable tool-transition patterns, which can be leveraged throughout the whole training process. Inspired by ant colony optimization where historically successful paths can be reflected by the pheromone, we propose Pheromone-Guided Policy Optimization (PhGPO), which learns a trajectory-based transition pattern (i.e., pheromone) from historical trajectories and then uses the learned pheromone to guide policy optimization. This learned pheromone provides explicit and reusable guidance that steers policy optimization toward historically successful tool transitions, thereby improving long-horizon tool planning. Comprehensive experimental results demonstrate the effectiveness of our proposed PhGPO.
Figures & tables
Figure 1: Two-stage pilot experiment with an explicit transition prior for long-horizon tool planning. We distill cross-trajectory transition memory from verified successful tool-use trajectories in an easier stage, then reuse it as a prior for GRPO optimization. Panels (a,b) show higher average return and success rate under matched interaction budgets, (c) fewer steps to the first successful tool-use trajectory, and (d) larger relative gains with longer trajectories.
Figure 2: Overview of PhGPO. PhGPO converts verified successful tool-use trajectories into a pheromone-based explicit transition prior over tool-transition and tool-to-invocation edges. The prior is updated through deposition and evaporation, and incorporated into rollout generation to guide the policy toward historically successful transitions and argument invocations. Training proceeds from supervised next-tool warm-up to progressive pheromone-guided reinforcement learning and finally full pheromone-guided policy optimization.
Method
Toolathlon
TRAJECT-Bench
TOUCAN
Qwen2.5-7B
Llama3.1-8B
Qwen2.5-7B
Llama3.1-8B
Qwen2.5-7B
Llama3.1-8B
Match R.
TSR
Match R.
TSR
Match R.
TSR
Match R.
TSR
Match R.
TSR
Match R.
TSR
Standard Prompting & Planning Strategies
ReAct
5.31
2.92
5.15
2.48
15.32
9.20
9.16
7.01
11.64
8.92
7.03
5.88
Plan-and-Solve
6.42
4.61
4.27
3.96
12.44
11.35
9.05
5.54
12.71
10.15
7.42
5.12
Beyond ReAct
3.85
3.10
5.23
4.18
11.33
10.74
15.61
14.20
12.92
10.84
7.35
6.54
Table 1: Main Results on Long-Horizon Tool-Use Benchmarks. We report Match Ratio (Match R., %) and Task Success Rate (TSR, %) on Toolathlon, TRAJECT-Bench, and TOUCAN across two backbone models. Higher values ( ↑ ) indicate better performance.
Figure 3: Emergence of the reference chain in pheromone transitions. As training proceeds, the transition matrix becomes increasingly concentrated on reference edges (red), while unrelated transitions fade due to evaporation, and pheromone values on the reference chain increase steadily.
Variant
Variant role
Match R. ↑
Drop M. ↓
TSR ↑
Drop TSR ↓
w/o Pheromone ( β=0 throughout)
removes explicit transition prior
21.23
4.02
16.72
3.62
w/o Mixed Curriculum
removes progressive oracle guidance
19.34
5.91
14.86
5.48
Static Prior
freezes pheromone updates
23.93
1.32
18.71
1.63
w/o Evaporation ( ρ=0 )
removes decay of stale statistics
21.72
3.53
16.28
4.06
w/o Task-dependent Pheromones
removes task-matched edge memory
22.45
2.80
17.35
2.99
PhGPO (Full)
all components enabled
25.25
–
20.34
–
Table 2: Ablation studies on Toolathlon. We report Match Ratio (Match R., %) and Task Success Rate (TSR, %) on the test set. Drop columns are computed relative to PhGPO (Full), where larger drops indicate stronger contribution of the removed component.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Tool source
L1 Tools
L2 Tools
#Instances
Avg Steps
Toolathlon
MCP apps
618
1,250
108 tasks
13.48
TRAJECT-Bench
executable APIs
381
715
5,870 queries
6.36
TOUCAN (Multi-step)
MCP servers
850
15,294
1,646,546 traj.
3.16
Appendix
Table 3: Dataset Statistics. We summarize the scale of the three benchmarks under our Tool-Transition Graph action representation. L1 Tools denotes the number of base tools, and L2 Tools denotes the number of argument-invocation patterns. Avg Steps reports the average number of tool calls per verifiably successful reference tool-use sequence in the retained episodes. For TOUCAN, we only consider Multi-step instances.
Figure 4: Ablation study of pheromone influence parameter β on Toolathlon benchmark. (a) Learning curves show that dynamic β annealing achieves the highest Match Ratio (25.25%), significantly outperforming fixed strategies. β=0 (no pheromone) exhibits slow, noisy convergence, while β=5 (over-guidance) becomes trapped in local optima. (b) Next-tool accuracy mirrors learning trends, with dynamic β reaching 27.16% through effective balance of exploration and exploitation. (c) Exploration diversity exhibits natural fluctuations due to finite-sample estimation from rollouts, but reveals clear trends: β=0 maintains excessively high diversity (unfocused exploration), β=5 shows extremely low diversity (trapped exploitation), while dynamic β demonstrates an optimal trajectory—starting with high exploration (0.75) and smoothly transitioning into the optimal range (0.55-0.65) that balances verified pattern reuse with adaptive exploration.
Setting
Match R.
TAcc
Final Diversity
β=0 (No pheromone)
21.23
19.86
0.76
β=1 (Fixed)
23.93
22.05
0.53
β=5 (Over-guidance)
17.80
16.50
0.25
Dynamic β (Ours)
25.25
27.16
0.58
Appendix
Table 4: Ablation on β for Toolathlon (Qwen2.5-7B). We report Match Ratio (Match R., %), Next-tool Accuracy (TAcc, %), and final Exploration Diversity.
Figure 5: Training Dynamics of PhGPO. (a) Average Return : All backbone models show a consistent upward trend in average return, indicating stable policy improvement during training. (b) Pheromone Graph Growth : The number of discovered edges grows rapidly in early epochs and stabilizes later, reflecting fast discovery of feasible tool transitions followed by refinement on a stable set of edges.
Figure 6: Hyperparameter sensitivity analysis on Toolathlon. We evaluate 8 hyperparameters for pheromone guidance and policy optimization. Red stars indicate our selected values; shaded regions show ±1 standard deviation across 3 seeds. Row 1: (a) βmax ; (b) ρ ; (c) α ; (d) w . Row 2: (e) M ; (f) η ; (g) ptf ; (h) γ .
Category
Parameter
Value
Description
Pheromone
βmax
0.7
Maximum pheromone influence weight in Eq. ( 7 )
ρ
0.01
Evaporation rate in Eq. ( 1 )
α
1.0
Deposition rate in Eq. ( 1 )
w
0.5
Task-dependent pheromone weight in Eq. ( 5 )
Optimization
M
5
GRPO group size, measured as rollouts per instance
η
10−7
Learning rate for RL fine-tuning
Appendix
Table 5: Key hyperparameters in PhGPO and selected values.
Optimizer variant
Qwen2.5-7B
Llama3.1-8B
Match R. ↑
TAcc ↑
Match R. ↑
TAcc ↑
PhGPO + GRPO
25.25
27.16
22.83
23.09
w/o pheromone guidance ( β=0 )
21.23
19.86
18.75
16.45
PhGPO + PPO
23.27
25.03
18.21
24.05
w/o pheromone guidance ( β=0 )
20.78
23.31
15.72
21.22
PhGPO + RLOO
20.21
24.47
18.51
22.69
Appendix
Table 6: PhGPO with different policy optimization backbones on Toolathlon. For each backbone optimizer, we report results with pheromone-guided sampling (Eq. ( 7 )) and without pheromone guidance by setting β=0 in sampling. Higher is better ( ↑ ).
Figure 7: Pheromone evolution on a longer tool-use trajectory, visualized on a 10-tool subset. Top: fused tool-transition pheromone τTool(⋅∣x) restricted to ten tools at checkpoint, with reference-chain edges highlighted in red. Bottom: pheromone values on the highlighted reference-chain transitions at the same checkpoint. The visualization shows that pheromone concentrates on reference-chain transitions while off-chain transitions fade under evaporation, indicating that the explicit transition prior remains selective on longer tool-use trajectories.
Figure 8: Time efficiency analysis on Toolathlon. We track the earliest training step at which a high-quality trajectory ( qτ(ξ)≥0.6 ) is first generated for each episode. PhGPO reaches broad coverage earlier than the ablation without pheromone, and the relative improvement remains broadly consistent across trajectory-length bins.
Method
Train step (s)
Train overhead
Peak GPU (GB)
Infer/step (ms)
Infer overhead
GRPO
11.86
baseline
17.3
58.5
baseline
+ graph maintenance
11.87
+0.01s
17.3
58.8
+0.3ms
+ graph + pheromone sampling
11.88
+0.02s
17.4
59.1
+0.6ms
Full PhGPO
11.89
+0.03s
17.4
59.3
+0.8ms
Appendix
Table 7: Runtime and memory overhead on Toolathlon. We report end-to-end overhead using Qwen2.5-7B.
Setting
β schedule
prand
Match R. ↑
TAcc ↑
Final Diversity
No pheromone
0
0.05
21.23
19.86
0.76
Full model w/o random exploration
dynamic β
0
24.63
26.34
0.52
Full model
dynamic β
0.05
25.25
27.16
0.58
Full model, larger random exploration
dynamic β
0.10
24.38
26.47
0.66
Appendix
Table 8: Random exploration and pheromone guidance on Toolathlon. We compare the effects of removing pheromone guidance and changing the random exploration probability prand using Qwen2.5-7B.