SkillGATE: Gate-Aware Monte Carlo Tree Search for Skill Retrieval
Authors: Rongchen Zhao, Yu Chen, Yanming Yang, Shijia Xu, Juyuan Wang, Jin Xu, Zibin Zheng, Jingping Liu
Organizations: School of Software Engineering, Sun Yat-Sen University, Guangdong, China · School of Future Technology, South China University of Technology, Guangdong, China · Institute for Math and AI, Wuhan University, Wuhan, China
Skill Retrieval (SR) aims to identify the most relevant skills from external skill libraries, and becomes increasingly challenging as libraries grow in scale and diversity. Existing methods either rank skills independently or rely on predefined graph propagation and hierarchical routing, making them vulnerable to semantic distractors, local trapping, and early routing errors. We formulate SR as an adaptive information-foraging process that coordinates region-level navigation with skill-level selection according to the utility and uncertainty observed during search. Based on this formulation, we propose SkillGATE, a graph-guided hierarchical retrieval framework with Gate-Aware Monte Carlo Tree Search (MCTS). SkillGATE constructs a graph-preserving hierarchical index and performs adaptive retrieval through selection, expansion, simulation, and backpropagation. G-PUCT guides action selection, expansion explores new regions, simulation evaluates candidate skills, and backpropagation updates search statistics. Experiments on six SR benchmarks show that SkillGATE consistently improves diverse retrieval and reranking backbones, achieving a 16.3% improvement in overall R@1 over the strongest retriever-based baseline. Our code is available at https://github.com/Edwinbe/SkillGATE-v1/.
Figures & tables
Figure 1: Comparison between different skill retrieval paradigms.
Figure 2: Framework Overview.
Method
TheoremQA
LogicBench
ToolQA
CHAMP
MedCalc.
BigCode.
Overall
R@1
R@10
R@1
R@10
R@1
R@10
R@1
R@10
R@1
R@10
R@1
R@10
R@1
R@10
Retriever-based Methods
BM25
57.2
80.7
12.0
36.1
7.0
55.1
13.2
36.1
29.3
69.2
23.6
61.1
23.0
59.3
BGE-Emb
66.8
86.1
4.1
20.5
32.2
83.4
9.8
34.0
41.4
70.1
20.7
62.1
31.6
65.7
Ada-002
68.0
90.8
9.1
31.6
18.8
54.8
13.7
41.7
82.6
98.4
18.2
52.0
36.9
64.2
GoS
54.8
83.4
7.6
25.5
6.8
20.4
4.9
25.1
67.1
81.8
17.2
49.3
28.0
48.6
Table 1: Main results (%) on six benchmarks for SR. Blue denotes retriever-based methods and purple denotes reranker-based methods. Bold marks the best in ours, underline marks the best baseline, and Improvement marks their performance gap.
Method
TheoremQA
LogicBench
ToolQA
CHAMP
MedCalc.
BigCode.
Overall
R@1
R@10
R@1
R@10
R@1
R@10
R@1
R@10
R@1
R@10
R@1
R@10
R@1
R@10
Ours
74.7
93.8
13.8
54.0
41.4
68.2
17.8
56.1
93.6
98.1
23.0
61.1
47.9
73.8
Indexing
w/ Graph G
58.9
74.6
15.3
51.3
8.8
14.1
15.4
50.4
70.4
80.5
10.8
25.8
29.9
45.3
w/ Tree T
24.2
29.2
4.5
11.7
4.7
5.6
7.8
21.6
36.1
36.9
4.3
7.9
13.8
17.3
Retrieval
Table 2: Ablation study on the indexing and retrieval components of our framework.
Hierarchy
Action Recovery
Prune Safety
Level
Nodes
Avg. Skills
Avg. Chid
Avg. Sim
Anode Fail
Askill Res.
Ahoriz Res.
Aprune Freq.
Aprune Acc.
L0
1
26,262.00
42.00
–
0.00%
0.00%
0.00%
0.00%
–
L1
42
625.29
8.67
0.22
30.56%
64.85%
35.15%
2.53%
95.49%
L2
338
77.52
4.30
0.32
15.19%
47.56%
64.63%
2.35%
95.46%
L3
1,126
21.17
3.21
0.41
10.16%
74.51%
58.82%
4.48%
94.97%
L4
880
12.01
0.00
0.43
20.22%
41.67%
47.22%
6.45%
96.19%
Table 3: Hierachy statics and level-wise analysis of actions.
Figure 3: Gate effectiveness (a) and skill-library (b) scalability.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Hyperparameter
Value
Index
Edge weights: sem./lex./label
0.35/0.20/0.45
Neighbor limit Knbr
50
Edge threshold τe
0.32
Depth lmax / Number nmin
4/20
Retrieval
Relevance weights: sem./lex./label
0.50/0.25/0.25
Gate Transition scale τκ
16
Appendix
Table 4: Hyperparameter settings.
Query-Relevance Weights
Reward Weights
Relevance Weights
Retrieval
Reward Weights
Retrieval
sem.
lex.
label
R@1
R@10
top-1
top- k1
top- k2
R@1
R@10
0.50
0.25
0.25
47.98±0.12
74.46±0.32
0.55
0.30
0.15
47.98±0.12
74.46±0.32
0.00
0.00
1.00
47.47±0.45
72.49±0.63
0.00
0.00
1.00
47.51±0.40
73.76±0.36
0.00
1.00
0.00
47.54±0.31
73.16±0.56
0.00
1.00
0.00
47.22±0.33
73.70±0.64
1.00
0.00
0.00
46.83±0.14
71.74±0.59
1.00
0.00
0.00
46.80±0.28
72.80±0.26
Appendix
Table 5: Sensitivity to query-relevance and reward weights.
Dataset
Query
Gold
R@1
R@10
BigCodeBench
114
311
23.92
63.08
CHAMP
22
36
12.88
41.67
LogicBench
76
76
14.47
56.58
MedCalcBench
110
110
93.64
97.27
TheoremQA
75
75
68.00
90.67
ToolQA
143
143
44.76
72.03
Appendix
Table 6: Sampled test set statistics.
Search Strategy
Hierarchy
Graph
Feedback
R@1
R@10
RW
×
✓
×
30.05
46.58
HKD
×
✓
×
30.52
47.06
PPR
×
✓
×
31.04
47.82
Greedy
✓
×
×
36.93
59.80
UCT
✓
✓
✓
42.16
65.51
PUCT
✓
✓
✓
43.73
67.62
Appendix
Table 7: Comparison of search strategies on the sampled subset.
Level
Vvisit Hit
Vhier Hit
Ours Hit
L0
100.0
100.0
100.0
L1
90.4
63.0
90.4
L2
64.0
39.2
66.8
L3
40.2
23.7
42.0
L4
28.0
16.3
32.2
Appendix
Table 8: Update-path comparison.
Ours
GoS
AgentSkillOS
Performance Metrics
R@1
47.88 (100.00%)
28.07 (58.63%)
24.30 (50.75%)
R@10
73.83 (100.00%)
48.69 (65.95%)
29.66 (40.17%)
Token Usage (/M)
LLM
56.08 (100.00%)
98.14 (175.00%)
848.21 (1512.50%)
Embedding
59.58 (100.00%)
38.85 (65.21%)
0.00 (0.00%)
Appendix
Table 9: Comparison of efficiency between Structured SR methods.
#
Phase
Operation
Scope Transition
Size
Gold?
1
Selection
Node-aware descent
root → p310
1214
No
2
Expansion
Skill-aware descent
p310 → p387
61
No
3
Simulation
Skill-aware descent
p387 → p22
48
No
4
Simulation
Node-aware descent
p22 → p24
17
No
5
Evaluation
Local candidate readouts
p24 → p24
17
No
Appendix
Table 10: Case study comparing the retrieval traces of Vanilla MCTS and SkillGATE .