Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a 4×11 grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.
Figures & tables
Figure 1: A deterministic two-stage retriever reaches observed parity with an LLM-mediated retrieval loop at zero in-loop LLM cost. Left: the task-solving loop (task → skill-retrieval module → agent → pytest verifier, which returns a graded reward in [0,1] ). Right: two methods that occupy the retrieval step, SkillSeek (top, bi-encoder + cross-encoder over MCP) and Liu et al. (2026b) ’s liu_refined (bottom, LLM agent refinement loop), with cost-vs-pass-rate Pareto plots on both pools.
192 curated skills
34K skill marketplace
Method
Family
Qwen3.5
MiniMax
Qwen3.5
MiniMax
none
reference
0.352
0.267
0.352
0.267
bm25
classical IR
0.430
0.387
0.420
0.346‡
bge-reranker-v2-m3
ours: OSS cross-encoder
0.480
0.371
0.409
0.346‡
Qwen3-Reranker-0.6B
ours: OSS LM-based reranker
0.430
0.353†
0.442‡
0.341
liu_hybrid
Liu et al. (2026b) , no refinement
0.470
0.378
0.417
0.335
Table 1: Main result: agent pass rate across pool, backbone, and method. Mean pass rate over 89 SkillsBench tasks per setting (missing trials counted as 0). Our family is tinted; bold marks the column winner. † 88/89: earthquake-phase-association timed out across most MiniMax-M2.7 conditions. ‡ Within-noise ties: 34K / MiniMax and 34K / Qwen3.5 (full numbers in § 4.4 ). Reranker-scaling and commercial-API methods appear only on the 34K pool with Qwen3.5 (§ 4.4 ).
Indexed text
pass rate
Δ
name + description
0.338
−2.2%
+ body[: 800 ]
0.328
−3.2%
+ Tool-REX v 3 (default)
0.360
0
Table 2: Indexing-text ablation. Agent pass rate on the 34K pool with Qwen3.5-397B-A17B under three choices for the text the bi-encoder and cross-encoder index. The default (Tool-REX v3) is strongest.
Stage-1 depth
pass rate
Δ
kinit=10
0.338
−2.1%
kinit=20 (default)
0.360
0
kinit=50
0.315
−4.5%
kinit=100
0.324
−3.5%
Table 3: Stage-1 candidate-depth ablation. Pass rate as we vary kinit , the number of candidates BGE-base passes to the cross-encoder. The default kinit=20 is the best choice.
Top- k to agent
pass rate
Δ
k=1
0.340
−2.0%
k=3
0.360
0
k=5 (default)
0.360
0
k=10
0.369
+0.9%
Table 4: Agent-side top- k ablation. Pass rate as we vary the number of reranked candidates returned via skill_lookup . Beyond k=3 , additional candidates buy nothing.
192 pool
34 K pool
Retrieval method
R@ 5
Δ
R@ 5
Δ
Stage 1 (bi-encoder retrieval, no cross-encoder)
bm25
0.546
n/a
0.391
n/a
BGE-base
0.546
±0.0%
0.379
−1.2%
Stage 2 ( + cross-encoder reranker)
+ bge-reranker-v2-m3
0.581
+3.5%
0.433
+5.4%
Table 5: First-stage recall ceiling. R@5 before and after our cross-encoder reranker on both pools. The reranker gain is much larger at the 34K scale because stage-1 recall is far from saturated there, whereas on the 192 pool it is already at ceiling.
Figure 2: Reranker scaling and OSS-vs-commercial. 34K pool with Qwen3.5-397B-A17B. (a) Pass rate plateaus within the Qwen3-Reranker family at the 0.6B size; Qwen3-Reranker-8B narrowly beats Voyage rerank-2.5 by +1.5%. (b) The helpfulness gap (pass-rate conditional on the gold skill being retrieved minus pass-rate conditional on a miss) explains why: Voyage rerank-2.5 is the only reranker with a negative helpfulness gap, meaning its retrievals do not translate into agent success.
Method
pass rate
Δ vs. none
retrieval cost
agent-loop cost
total
none
0.352
n/a
n/a
$27.41
$27.41
ours (BGE + bge-rrk-v2-m3)
0.409
+5.7%
$0 (CPU)
$27.54
$27.54
liu_hybrid
0.417
+6.5%
$0
$39.43
$39.43
liu_refined
0.442
+9.0%
$25.63
$25.67
$51.30
Table 6: Cost breakdown. 34K pool with Qwen3.5-397B-A17B. Retrieval cost is the pre-agent LLM spend (refinement); agent-loop cost is the main agent’s own token spend on the same OpenHands SDK harness. Our deterministic retriever pays nothing on the retrieval side and leaves the agent-loop cost essentially unchanged from the no-skill baseline.
Task
Δ pass rate
Top wins (rerank surfaces gold)
flood-risk-analysis
+1.00
pg-essay-to-audiobook
+1.00
gravitational-wave-detection
+0.89
earthquake-plate-calculation
+0.88
grid-dispatch-operator
+0.83
Table 7: Per-task wins and losses. Rerank vs. none on the 34K pool with MiniMax-M2.7; top-5 in each direction.
Difficulty
n
none
rerank
lift
easy
6
0.254
0.306
+5.2%
medium
52
0.324
0.378
+5.4%
hard
26
0.194
0.274
+8.0%
Table 8: Per-difficulty lift. Rerank-vs- none lift on the 34K pool with MiniMax-M2.7; difficulty labels come from task.toml , with 3 of 89 tasks unlabeled.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Short name (paper)
Canonical identifier
Host
BGE-base
BAAI/bge-base-en-v1.5
HuggingFace
bge-reranker-v2-m3
BAAI/bge-reranker-v2-m3
HuggingFace
Qwen3-Reranker-0.6B
Qwen/Qwen3-Reranker-0.6B
HuggingFace
Qwen3-Reranker-4B
Qwen/Qwen3-Reranker-4B
HuggingFace
Qwen3-Reranker-8B
Qwen/Qwen3-Reranker-8B
HuggingFace
Qwen3-Embedding-4B
Qwen/Qwen3-Embedding-4B
HuggingFace
Appendix
Table 9: Canonical model identifiers. Used in this paper across the main grid (§ 4.2 ), analysis (§ 4.4 ), and appendix.
Condition
pass rate
tokens/trial
none (no skills loaded)
0.384
268 K
all (entire 192-skill pool loaded)
0.387
426 K (+59%)
oracle (per-task curated subset, 1–3 skills)
0.484
314 K (+17%)
Appendix
Table 10: Loading every skill from a small pool already fails. On 89 SkillsBench tasks with a locally served gpt-oss-120b backbone, loading the full 192-skill pool ( all ) yields the same pass rate as loading nothing (0.387 vs. 0.384) despite a 59% token surcharge, while a per-task curated subset ( oracle ) lifts pass rate to 0.484.
Backbone
OpenRouter spend (USD)
Qwen3.5-397B-A17B (primary)
617.77
MiniMax-M2.7 (secondary)
149.19
GLM 5 (exploratory)
45.84
Kimi K2.6 (exploratory)
24.67
Total
837.47
Appendix
Table 11: OpenRouter API spend per backbone. Totals for the four LLM backbones queried during this work. The two locally served models (gpt-oss-120b for the Tool-REX v3 index expansion and Qwen3-Embedding-4B for Liu et al. (2026b) ’s baseline retriever) ran on the two-GPU vLLM stack described above and are not on OpenRouter. Captured from the OpenRouter usage dashboard on 2026-05-24.
As LLM agents are increasingly deployed with large libraries of reusable skills, selecting the right skill for a user request has become a critical systems challenge. In small libraries, users may invoke skills explicitly by name, but this assumption breaks down as skill ecosystems grow under tight context and latency budgets. Despite its practical importance, skill retrieval remains underexplored, with limited benchmarks and little understanding of retrieval behavior on realistic skill libraries. To address this gap, we introduce SkillRet, a large-scale benchmark for skill retrieval in LLM agents. SkillRet contains 16,129 public agent skills, organized with structured semantic tags and a two-level taxonomy spanning 6 major categories and 18 sub-categories. It provides 63,259 training samples and 4,392 evaluation queries with disjoint skill pools, enabling both benchmarking and retrieval-oriented training. Across a diverse set of retrievers, we find that skill retrieval remains far from solved: off-the-shelf models struggle on realistic large-scale skill libraries, and prior skill-retrieval models still leave substantial headroom. Task-specific fine-tuning on SkillRet improves NDCG@10 by 12.9 points over the strongest prior retriever and by 16.2 points over the strongest off-the-shelf retriever. Our analysis further suggests that these gains arise because fine-tuned models better focus on the small skill-relevant signals within long and noisy queries. These results establish SkillRet as a strong benchmark and foundation for future research on retrieval in large-scale agent systems. We publicly release the benchmark (https://huggingface.co/datasets/ThakiCloud/SKILLRET), code (https://github.com/ThakiCloud/SKILLRET), and model checkpoints (0.6B: https://huggingface.co/ThakiCloud/SKILLRET-Embedding-0.6B; 8B: https://huggingface.co/ThakiCloud/SKILLRET-Embedding-8B).
As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution. Existing retrievers often treat skill contents as ordinary documents, overlooking their highly regular structure: shared descriptive patterns recur across many skills while providing little evidence for distinguishing the required capability. We show that this shared descriptive background is reflected in dense relevance scores, induces a pronounced energy gap between queries and skill documents, and obscures discriminative signals, especially for structurally similar hard negatives. Based on this observation, we propose SkillSight, a training-free retrieval framework that calibrates shared background in both semantic and lexical spaces. Semantic Background Calibration estimates a background subspace from generic tokens identified by IDF, reducing similarity induced by shared descriptive patterns, while Lexical Evidence Calibration downweights shared background tokens to recover discriminative token-level evidence. Experiments on SRA-Bench and SkillBench-Supp demonstrate consistent improvements across retrieval metrics, with SkillSight improving Recall@10 by up to 20.21 percentage points over the original dense retriever. It is up to 1,248 times faster than the Dense + Reranker baseline. In end-to-end evaluation, SkillSight achieves the best overall performance across three agent models and outperforms LLM Selection by up to 4.97 percentage points. These results identify shared descriptive background as a source of ranking interference in skill retrieval and demonstrate that calibrating it enables accurate and efficient skill selection without additional training. Our code can be found at https://github.com/xiaojinying/SkillSight
Jinying Xiao, Bin Li, Xiaopeng Li +7
1National University of Defense Technology · 2Qinghai Normal University · 3Xizang University
Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit. Existing benchmarks provide limited cov- erage of such requests. To address this gap, we introduce SkillReason-Bench, a large-scale cross-domain benchmark containing 3,729 queries and a retrieval corpus of 61,228 skills spanning nine domains. We further propose SkillRea- son, a two-stage framework that uses chain-of-thought rea- soning as training-time supervision for skill retrieval. In Stage I, capability reasoning traces generated by a stronger teacher provide explicit supervision through contrastive learning, re- trieval distribution alignment, and language modeling, en- couraging the retriever to internalize capability reasoning in its query representation. In Stage II, a retrieval-guided GRPO objective encourages the model to explore reasoning trajecto- ries better suited to its own capabilities and more effective for retrieval. At inference, SkillReason directly encodes the orig- inal query without autoregressive CoT generation, preserv- ing efficient query-only retrieval. Extensive experiments on SkillReason-Bench, SkillRet, and SRA-Bench show that Skill- Reason achieves state-of-the-art performance across all three benchmarks, demonstrating that reasoning-enhanced training better bridges the semantic gap between high-level task goals and skill capabilities.
Donghong Jiang, Endian Lin, Luoping Cui +6
Beijing University of Posts and Telecommunications · Peking University · Beijing ZOYEN Technology Co., Ltd. +1