Similar agent skills can share instructions but differ in their conditions of use. Query-based text selection may retain shared instructions and omit these distinctions. We introduce SkillContrast, a training-free selector that compares retrieved skills and retains their differing text with local context for a pretrained reranker. On 1,235 requests from SameCapRisk-Bench, it yields 54-72 more clean hits (requests that retrieve a helpful skill without its marked risky sibling) than TF-IDF query selection at identical per-candidate input lengths, across 2 retrievers and 2 reranker sizes. Length-matched component replacements identify differing text as the main contributor in the primary setting, with smaller, mixed context effects. Relative to full skill bodies, SkillContrast uses 51.1-58.8% fewer model-input tokens, with 10-18 fewer clean hits at 0.6B and matching or higher observed clean-hit counts at 4B. Candidate-relative differences thus complement query relevance in selecting compact reranking inputs.
Figures & tables
Figure 1. SkillContrast in a reranking pipeline. Our text-selection stage (“Select passages”) retains differences (teal), preceding context (outlined), and opening context. The frozen reranker and MMR are existing components. Full-body similarity feeds MMR separately; the query-based no-neighbor branch is not shown. The query branches into passage selection and pointwise reranking. Candidate skills feed passage selection and a full-body-similarity path. Relevance ranks and similarity enter greedy selection, which returns 3 candidates.
Generic retrieval controls
Public skill retrieval methods
Ours
BGE-M3
BGE- reranker
Qwen- listwise
RRF
SkillRouter
SkillRet
R3-Skill
SkillSight
SkillContrast
R@3 ↑
0.718
0.750
0.764
0.744
0.829
0.801
0.846
0.837
0.867
HSR@3 ↓
0.172
0.155
0.163
0.228
0.165
0.150
0.137
0.165
0.113
CH@3 ↑
0.657
0.721
0.704
0.632
0.777
0.758
0.813
0.767
0.837
Table 1. Retrieval comparison on 1,235 requests. All methods use MMR(0.75) on their own top-20 candidates; SkillContrast reranks SkillSight candidates with 0.6B. Bold: best overall; underlined: best reference.
SkillSight / 0.6B
SkillSight / 4B
SkillRet / 0.6B
SkillRet / 4B
Selection strategy
R ↑
HSR ↓
CH ↑
R ↑
HSR ↓
CH ↑
R ↑
HSR ↓
CH ↑
R ↑
HSR ↓
CH ↑
Full body
0.888
0.105
0.852
0.915
0.073
0.898
0.870
0.100
0.837
0.901
0.063
0.886
Prefix
0.828
0.156
0.784
0.840
0.135
0.821
0.791
0.165
0.753
0.811
0.134
0.796
Query-selected
0.829
0.151
0.794
0.877
0.103
0.860
0.803
0.155
0.773
0.842
0.113
0.828
SkillContrast (ours)
0.867
0.113
0.837
0.923
0.061
0.906
0.856
0.099
0.829
0.900
0.057
0.886
Query-context variant
0.871
0.113
0.840
0.922
0.062
0.908
0.854
0.105
0.826
0.899
0.061
0.884
Table 2. Text selection on 1,235 requests. Each group fixes candidates, reranker, and MMR(0.75); compact strategies match length. All metrics use K=3 . Bold: best, including ties. Token row: rounded means per candidate, full → compact; reductions use exact totals.
Input variant
Helpful ↑
Risky ↓
Clean ↑
Δ Clean
SkillContrast (ours)
1,071
139
1,034
—
Replace differences
995
203
956
−78
Replace first block
1,074
133
1,034
0
Replace preceding context
1,076
139
1,038
+4
BM25-selected
1,000
225
944
−90
Table 3. Component contributions at matched input length (SkillSight, 0.6B, MMR(0.75), 1,235 requests). Top-3 hit counts; Δ Clean is relative to SkillContrast (dash: reference). BM25 is a separate selector control.
Figure 2. Two migration skills, 1 requested target (SkillSight, 0.6B). Excerpts are abridged; inputs match length per candidate (A: 328 tokens; B: 327). Final ranks use the same MMR over 20 candidates. Match labels are offline interpretations. A matrix aligns 2 skills with 2 passage selectors. The request targets Java 21 and Spring Boot 3.2. Both query-selected excerpts refer to a configured target; SkillContrast exposes Java 21 and Spring Boot 3.2 for A, versus Java 8 and Spring Boot 2.7 for B. A moves from final rank 5 to 1, and B from 1 to 5.