Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a 4×11 grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.
Figures & tables
Figure 1: A deterministic two-stage retriever reaches observed parity with an LLM-mediated retrieval loop at zero in-loop LLM cost. Left: the task-solving loop (task → skill-retrieval module → agent → pytest verifier, which returns a graded reward in [0,1] ). Right: two methods that occupy the retrieval step, SkillSeek (top, bi-encoder + cross-encoder over MCP) and Liu et al. (2026b) ’s liu_refined (bottom, LLM agent refinement loop), with cost-vs-pass-rate Pareto plots on both pools.
192 curated skills
34K skill marketplace
Method
Family
Qwen3.5
MiniMax
Qwen3.5
MiniMax
none
reference
0.352
0.267
0.352
0.267
bm25
classical IR
0.430
0.387
0.420
0.346‡
bge-reranker-v2-m3
ours: OSS cross-encoder
0.480
0.371
0.409
0.346‡
Qwen3-Reranker-0.6B
ours: OSS LM-based reranker
0.430
0.353†
0.442‡
0.341
liu_hybrid
Liu et al. (2026b) , no refinement
0.470
0.378
0.417
0.335
Table 1: Main result: agent pass rate across pool, backbone, and method. Mean pass rate over 89 SkillsBench tasks per setting (missing trials counted as 0). Our family is tinted; bold marks the column winner. † 88/89: earthquake-phase-association timed out across most MiniMax-M2.7 conditions. ‡ Within-noise ties: 34K / MiniMax and 34K / Qwen3.5 (full numbers in § 4.4 ). Reranker-scaling and commercial-API methods appear only on the 34K pool with Qwen3.5 (§ 4.4 ).
Indexed text
pass rate
Δ
name + description
0.338
−2.2%
+ body[: 800 ]
0.328
−3.2%
+ Tool-REX v 3 (default)
0.360
0
Table 2: Indexing-text ablation. Agent pass rate on the 34K pool with Qwen3.5-397B-A17B under three choices for the text the bi-encoder and cross-encoder index. The default (Tool-REX v3) is strongest.
Stage-1 depth
pass rate
Δ
kinit=10
0.338
−2.1%
kinit=20 (default)
0.360
0
kinit=50
0.315
−4.5%
kinit=100
0.324
−3.5%
Table 3: Stage-1 candidate-depth ablation. Pass rate as we vary kinit , the number of candidates BGE-base passes to the cross-encoder. The default kinit=20 is the best choice.
Top- k to agent
pass rate
Δ
k=1
0.340
−2.0%
k=3
0.360
0
k=5 (default)
0.360
0
k=10
0.369
+0.9%
Table 4: Agent-side top- k ablation. Pass rate as we vary the number of reranked candidates returned via skill_lookup . Beyond k=3 , additional candidates buy nothing.
192 pool
34 K pool
Retrieval method
R@ 5
Δ
R@ 5
Δ
Stage 1 (bi-encoder retrieval, no cross-encoder)
bm25
0.546
n/a
0.391
n/a
BGE-base
0.546
±0.0%
0.379
−1.2%
Stage 2 ( + cross-encoder reranker)
+ bge-reranker-v2-m3
0.581
+3.5%
0.433
+5.4%
Table 5: First-stage recall ceiling. R@5 before and after our cross-encoder reranker on both pools. The reranker gain is much larger at the 34K scale because stage-1 recall is far from saturated there, whereas on the 192 pool it is already at ceiling.
Figure 2: Reranker scaling and OSS-vs-commercial. 34K pool with Qwen3.5-397B-A17B. (a) Pass rate plateaus within the Qwen3-Reranker family at the 0.6B size; Qwen3-Reranker-8B narrowly beats Voyage rerank-2.5 by +1.5%. (b) The helpfulness gap (pass-rate conditional on the gold skill being retrieved minus pass-rate conditional on a miss) explains why: Voyage rerank-2.5 is the only reranker with a negative helpfulness gap, meaning its retrievals do not translate into agent success.
Method
pass rate
Δ vs. none
retrieval cost
agent-loop cost
total
none
0.352
n/a
n/a
$27.41
$27.41
ours (BGE + bge-rrk-v2-m3)
0.409
+5.7%
$0 (CPU)
$27.54
$27.54
liu_hybrid
0.417
+6.5%
$0
$39.43
$39.43
liu_refined
0.442
+9.0%
$25.63
$25.67
$51.30
Table 6: Cost breakdown. 34K pool with Qwen3.5-397B-A17B. Retrieval cost is the pre-agent LLM spend (refinement); agent-loop cost is the main agent’s own token spend on the same OpenHands SDK harness. Our deterministic retriever pays nothing on the retrieval side and leaves the agent-loop cost essentially unchanged from the no-skill baseline.
Task
Δ pass rate
Top wins (rerank surfaces gold)
flood-risk-analysis
+1.00
pg-essay-to-audiobook
+1.00
gravitational-wave-detection
+0.89
earthquake-plate-calculation
+0.88
grid-dispatch-operator
+0.83
Table 7: Per-task wins and losses. Rerank vs. none on the 34K pool with MiniMax-M2.7; top-5 in each direction.
Difficulty
n
none
rerank
lift
easy
6
0.254
0.306
+5.2%
medium
52
0.324
0.378
+5.4%
hard
26
0.194
0.274
+8.0%
Table 8: Per-difficulty lift. Rerank-vs- none lift on the 34K pool with MiniMax-M2.7; difficulty labels come from task.toml , with 3 of 89 tasks unlabeled.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Short name (paper)
Canonical identifier
Host
BGE-base
BAAI/bge-base-en-v1.5
HuggingFace
bge-reranker-v2-m3
BAAI/bge-reranker-v2-m3
HuggingFace
Qwen3-Reranker-0.6B
Qwen/Qwen3-Reranker-0.6B
HuggingFace
Qwen3-Reranker-4B
Qwen/Qwen3-Reranker-4B
HuggingFace
Qwen3-Reranker-8B
Qwen/Qwen3-Reranker-8B
HuggingFace
Qwen3-Embedding-4B
Qwen/Qwen3-Embedding-4B
HuggingFace
Appendix
Table 9: Canonical model identifiers. Used in this paper across the main grid (§ 4.2 ), analysis (§ 4.4 ), and appendix.
Condition
pass rate
tokens/trial
none (no skills loaded)
0.384
268 K
all (entire 192-skill pool loaded)
0.387
426 K (+59%)
oracle (per-task curated subset, 1–3 skills)
0.484
314 K (+17%)
Appendix
Table 10: Loading every skill from a small pool already fails. On 89 SkillsBench tasks with a locally served gpt-oss-120b backbone, loading the full 192-skill pool ( all ) yields the same pass rate as loading nothing (0.387 vs. 0.384) despite a 59% token surcharge, while a per-task curated subset ( oracle ) lifts pass rate to 0.484.
Backbone
OpenRouter spend (USD)
Qwen3.5-397B-A17B (primary)
617.77
MiniMax-M2.7 (secondary)
149.19
GLM 5 (exploratory)
45.84
Kimi K2.6 (exploratory)
24.67
Total
837.47
Appendix
Table 11: OpenRouter API spend per backbone. Totals for the four LLM backbones queried during this work. The two locally served models (gpt-oss-120b for the Tool-REX v3 index expansion and Qwen3-Embedding-4B for Liu et al. (2026b) ’s baseline retriever) ran on the two-GPU vLLM stack described above and are not on OpenRouter. Captured from the OpenRouter usage dashboard on 2026-05-24.