Similar agent skills can share instructions but differ in their conditions of use. Query-based text selection may retain shared instructions and omit these distinctions. We introduce SkillContrast, a training-free selector that compares retrieved skills and retains their differing text with local context for a pretrained reranker. On 1,235 requests from SameCapRisk-Bench, it yields 54-72 more clean hits (requests that retrieve a helpful skill without its marked risky sibling) than TF-IDF query selection at identical per-candidate input lengths, across 2 retrievers and 2 reranker sizes. Length-matched component replacements identify differing text as the main contributor in the primary setting, with smaller, mixed context effects. Relative to full skill bodies, SkillContrast uses 51.1-58.8% fewer model-input tokens, with 10-18 fewer clean hits at 0.6B and matching or higher observed clean-hit counts at 4B. Candidate-relative differences thus complement query relevance in selecting compact reranking inputs.
Figures & tables
Figure 1. SkillContrast in a reranking pipeline. Our text-selection stage (“Select passages”) retains differences (teal), preceding context (outlined), and opening context. The frozen reranker and MMR are existing components. Full-body similarity feeds MMR separately; the query-based no-neighbor branch is not shown. The query branches into passage selection and pointwise reranking. Candidate skills feed passage selection and a full-body-similarity path. Relevance ranks and similarity enter greedy selection, which returns 3 candidates.
Generic retrieval controls
Public skill retrieval methods
Ours
BGE-M3
BGE- reranker
Qwen- listwise
RRF
SkillRouter
SkillRet
R3-Skill
SkillSight
SkillContrast
R@3 ↑
0.718
0.750
0.764
0.744
0.829
0.801
0.846
0.837
0.867
HSR@3 ↓
0.172
0.155
0.163
0.228
0.165
0.150
0.137
0.165
0.113
CH@3 ↑
0.657
0.721
0.704
0.632
0.777
0.758
0.813
0.767
0.837
Table 1. Retrieval comparison on 1,235 requests. All methods use MMR(0.75) on their own top-20 candidates; SkillContrast reranks SkillSight candidates with 0.6B. Bold: best overall; underlined: best reference.
SkillSight / 0.6B
SkillSight / 4B
SkillRet / 0.6B
SkillRet / 4B
Selection strategy
R ↑
HSR ↓
CH ↑
R ↑
HSR ↓
CH ↑
R ↑
HSR ↓
CH ↑
R ↑
HSR ↓
CH ↑
Full body
0.888
0.105
0.852
0.915
0.073
0.898
0.870
0.100
0.837
0.901
0.063
0.886
Prefix
0.828
0.156
0.784
0.840
0.135
0.821
0.791
0.165
0.753
0.811
0.134
0.796
Query-selected
0.829
0.151
0.794
0.877
0.103
0.860
0.803
0.155
0.773
0.842
0.113
0.828
SkillContrast (ours)
0.867
0.113
0.837
0.923
0.061
0.906
0.856
0.099
0.829
0.900
0.057
0.886
Query-context variant
0.871
0.113
0.840
0.922
0.062
0.908
0.854
0.105
0.826
0.899
0.061
0.884
Table 2. Text selection on 1,235 requests. Each group fixes candidates, reranker, and MMR(0.75); compact strategies match length. All metrics use K=3 . Bold: best, including ties. Token row: rounded means per candidate, full → compact; reductions use exact totals.
Input variant
Helpful ↑
Risky ↓
Clean ↑
Δ Clean
SkillContrast (ours)
1,071
139
1,034
—
Replace differences
995
203
956
−78
Replace first block
1,074
133
1,034
0
Replace preceding context
1,076
139
1,038
+4
BM25-selected
1,000
225
944
−90
Table 3. Component contributions at matched input length (SkillSight, 0.6B, MMR(0.75), 1,235 requests). Top-3 hit counts; Δ Clean is relative to SkillContrast (dash: reference). BM25 is a separate selector control.
Figure 2. Two migration skills, 1 requested target (SkillSight, 0.6B). Excerpts are abridged; inputs match length per candidate (A: 328 tokens; B: 327). Final ranks use the same MMR over 20 candidates. Match labels are offline interpretations. A matrix aligns 2 skills with 2 passage selectors. The request targets Java 21 and Spring Boot 3.2. Both query-selected excerpts refer to a configured target; SkillContrast exposes Java 21 and Spring Boot 3.2 for A, versus Java 8 and Spring Boot 2.7 for B. A moves from final rank 5 to 1, and B from 1 to 5.
As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution. Existing retrievers often treat skill contents as ordinary documents, overlooking their highly regular structure: shared descriptive patterns recur across many skills while providing little evidence for distinguishing the required capability. We show that this shared descriptive background is reflected in dense relevance scores, induces a pronounced energy gap between queries and skill documents, and obscures discriminative signals, especially for structurally similar hard negatives. Based on this observation, we propose SkillSight, a training-free retrieval framework that calibrates shared background in both semantic and lexical spaces. Semantic Background Calibration estimates a background subspace from generic tokens identified by IDF, reducing similarity induced by shared descriptive patterns, while Lexical Evidence Calibration downweights shared background tokens to recover discriminative token-level evidence. Experiments on SRA-Bench and SkillBench-Supp demonstrate consistent improvements across retrieval metrics, with SkillSight improving Recall@10 by up to 20.21 percentage points over the original dense retriever. It is up to 1,248 times faster than the Dense + Reranker baseline. In end-to-end evaluation, SkillSight achieves the best overall performance across three agent models and outperforms LLM Selection by up to 4.97 percentage points. These results identify shared descriptive background as a source of ranking interference in skill retrieval and demonstrate that calibrating it enables accurate and efficient skill selection without additional training. Our code can be found at https://github.com/xiaojinying/SkillSight
Jinying Xiao, Bin Li, Xiaopeng Li +7
1National University of Defense Technology · 2Qinghai Normal University · 3Xizang University
Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a 4×11 grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.
Guanqun Yang, Wenlong Zhang, Tian Shi +1
Stevens Institute of Technology, Hoboken, NJ, USA · Independent Researcher
Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retrieving the appropriate skill from a large- scale library remains challenging because realistic user re- quests are often concise and underspecified, stating only the task goal while leaving the required capabilities and execu- tion steps implicit. Existing benchmarks provide limited cov- erage of such requests. To address this gap, we introduce SkillReason-Bench, a large-scale cross-domain benchmark containing 3,729 queries and a retrieval corpus of 61,228 skills spanning nine domains. We further propose SkillRea- son, a two-stage framework that uses chain-of-thought rea- soning as training-time supervision for skill retrieval. In Stage I, capability reasoning traces generated by a stronger teacher provide explicit supervision through contrastive learning, re- trieval distribution alignment, and language modeling, en- couraging the retriever to internalize capability reasoning in its query representation. In Stage II, a retrieval-guided GRPO objective encourages the model to explore reasoning trajecto- ries better suited to its own capabilities and more effective for retrieval. At inference, SkillReason directly encodes the orig- inal query without autoregressive CoT generation, preserv- ing efficient query-only retrieval. Extensive experiments on SkillReason-Bench, SkillRet, and SRA-Bench show that Skill- Reason achieves state-of-the-art performance across all three benchmarks, demonstrating that reasoning-enhanced training better bridges the semantic gap between high-level task goals and skill capabilities.
Donghong Jiang, Endian Lin, Luoping Cui +6
Beijing University of Posts and Telecommunications · Peking University · Beijing ZOYEN Technology Co., Ltd. +1