Research
SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale
Overview Research area: Information retrieval (cs.IR) applied to LLM agent tooling — specifically, how an agent should select reusable procedural skills from a large skill marketplace at task time. Te

- arXiv
- 2609.38822
- Published
- 2026-09-30
- Authors
- Guanqun Yang, Wenlong Zhang, Tian Shi, Ping Wang
AI summary
Overview
Research area: Information retrieval (cs.IR) applied to LLM agent tooling — specifically, how an agent should select reusable procedural skills from a large skill marketplace at task time.
Technical level: Advanced. The paper assumes familiarity with two-stage retrieval architectures (bi-encoders, cross-encoders), BM25, reciprocal-rank fusion, recall metrics like R@5, and agent harnesses that speak the Model Context Protocol (MCP).
One-sentence scope: The paper benchmarks a deterministic, training-free two-stage retrieval recipe against an LLM-mediated agentic retrieval loop across a 4 × 11 grid of skill pool, agent backbone, and retrieval method on the 89-task SkillsBench benchmark, and finds the deterministic recipe reaches observed pass-rate parity at zero in-loop LLM cost.
What This Paper Is About
Anthropic's Agent Skills package reusable procedural knowledge into SKILL.md directories, and public aggregations of these skills have grown past 230,000, which makes choosing the right skill harder than authoring one. The dominant approach in the literature outsources this choice to the LLM agent itself: the agent rewrites queries and refines candidate skills inside its own decision loop, spending LLM tokens on every task. SkillSeek asks whether a conventional, non-LLM information-retrieval pipeline can match that approach at a fraction of the cost.
Key Contributions
-
SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe — a BGE-base bi-encoder (110M parameters, 768-dim embeddings) retrieving top-20 candidates, feeding a small cross-encoder reranker that returns top-5 — exposed as an MCP server so any MCP-capable harness can adopt it without source-level changes.
-
An empirical head-to-head comparison at marketplace scale across a 4 × 11 grid of (pool, backbone, method) on the 89-task SkillsBench benchmark, contrasting deterministic retrieval (
none,bm25, SkillSeek variants) against two reimplementations of Liu et al. (2026b)'s LLM-mediated loop (liu_hybridandliu_refined). -
A cost accounting decomposing per-trial spend into a retrieval-side component and a main-agent-loop component, showing total per-trial spend falls from $51.30 to $27.54 while pass rate stays comparable.
-
A mechanistic explanation and diagnostics — a first-stage recall-ceiling account of why lightweight rerankers suffice on curated pools but heavier ones help on marketplace pools, plus a proposed "helpfulness gap" diagnostic that separates retrieval recall from actual agent usefulness.
-
Full reproducibility artifacts, including scripts that regenerate every number in the paper, released at the linked GitHub repository.
Main Findings
-
Deterministic retrieval reaches observed parity with the LLM-mediated loop: Plain
bm25alone records a pass rate at or aboveliu_refinedon three of the four (pool, backbone) settings: 192/Qwen3.5 (0.430 vs. 0.397), 192/MiniMax (0.387 vs. 0.344), and 34K/MiniMax (0.346 vs. 0.329). Only on 34K/Qwen3.5 doesliu_refinedleadbm25(0.442 vs. 0.420), and there SkillSeek's Qwen3-Reranker-0.6B variant ties it at 0.442. -
The best cross-encoder is at or above
liu_refinedon every other setting, with the largest observed difference on 192/Qwen3.5 (0.480 vs. 0.397); on 34K/MiniMax the bge-reranker-v2-m3 variant (0.346) equalsbm25. The authors explicitly frame these as observed differences on a single trial per cell, not a resolved advantage. -
Which method wins depends on pool size: On the small curated pool the simplest methods lead (
bm25on MiniMax-M2.7, bge-reranker-v2-m3 on Qwen3.5-397B-A17B). On the 34K marketplace pool the heavier methods catch up (Qwen3-Reranker-0.6B tiesliu_refinedat 0.442). Tracking across pools on Qwen3.5-397B-A17B, bge-reranker-v2-m3 drops 7.1% andliu_hybriddrops 5.3%, while Qwen3-Reranker-0.6B rises 1.2% andliu_refinedrises 4.5%. -
The same split appears inside Liu et al.'s own method: On the 192 pool,
liu_hybridbeatsliu_refinedby +7.3% on Qwen3.5-397B-A17B (0.470 vs. 0.397) and +3.4% on MiniMax-M2.7 (0.378 vs. 0.344); the order reverses on the 34K Qwen3.5 pool, whereliu_refinedbeatsliu_hybridby +2.5% (0.442 vs. 0.417). -
Loading the whole pool does not work: Even at modest size (192 skills), loading every skill into the agent's context leaves pass rate at the no-skill baseline while inflating token cost by 59%. Curated subsets are also non-monotonic: one relevant skill lifts the agent by +17.8%, two-to-three by +18.6%, but four or more regress to +5.9%.
-
Cost drops sharply: SkillSeek's two models run on CPU with zero LLM tokens, leaving agent spend within fifty cents of the no-skill floor ($27.54 vs. $27.41). By contrast
liu_hybridinflates the agent loop by 44% ($39.43 vs. $27.41), andliu_refinedtotals $51.30 for +9.0% over no skills. Median latency is approximately 1.1 seconds perskill_lookupcall (1.9 seconds at p95). -
Indexing text dominates the design space: Replacing the Tool-REX v3 four-tag profile with plain name+description drops pass rate by 2.2% on 34K/Qwen3.5; appending the raw
SKILL.mdbody drops it a further 3.2%. -
Stage-1 depth of 20 is optimal: Halving to 10 costs 2.1%; doubling to 50 costs 4.5% and quintupling to 100 costs 3.5%. Deeper candidate lists perturb the rank-1 choice rather than help it.
-
Top-3 captures essentially all the gain:
k=1costs 2.0% relative tok=5, whilek=3,k=5, andk=10are tied within 1%. Beyondk=3, additional candidates buy nothing. -
Reranker scaling plateaus near 0.6B: Within the Qwen3-Reranker family on 34K/Qwen3.5, 0.6B reaches 0.442, 4B drops to 0.430, and 8B recovers to 0.450 — a marginal +0.7% over 0.6B from a 13× parameter increase. The strongest single-reranker configuration only matches
liu_refined(0.442). -
Open-source rerankers beat the commercial alternatives tested: Qwen3-Reranker-8B (0.450) narrowly beats Voyage rerank-2.5 (0.435) by +1.5%, with 13 wins, 67 ties, and 9 losses on the 89-task paired comparison.
-
A "helpfulness gap" explains why high recall is not enough: Voyage rerank-2.5 has the highest top-5 recall (72.2%) but a negative helpfulness gap (Δ = −0.025), meaning its retrievals do not translate into agent success; Qwen3-Reranker-8B sits at Δ = +0.149 and bge-reranker-v2-m3 at Δ = +0.012. The mechanism is that the agent commits to the rank-1 candidate on roughly 65% of tasks, so rank-1 usefulness matters more than whether the gold skill appears anywhere in top-5.
-
Per-task wins and losses: On 34K/MiniMax, the largest positive lifts (+0.89 to +1.00) occur where the gold skill's name and description closely match the query (for example
flood-risk-analysis,pg-essay-to-audiobook,gravitational-wave-detection). The largest negative lifts (−0.50 to −0.80, e.g.parallel-tfidf-search,energy-ac-optimal-power-flow) occur where another pool skill is a near-twin in description — ranking failures, not coverage failures. -
Hard tasks benefit most: Stratifying rerank-vs-none lift by SkillsBench difficulty on 34K/MiniMax gives +8.0% on hard tasks (n=26, baseline 0.194), +5.4% on medium (n=52, baseline 0.324), and +5.2% on easy (n=6, baseline 0.254).
-
Recall ceiling evidence: On the 192 pool,
bm25alone reaches R@5 = 0.546 and BGE-base matches it exactly, with the cross-encoder adding only +3.5%. On the 34K pool,bm25is at 0.391 and BGE-base at 0.379, with the cross-encoder adding +5.4%, growing to +9.9% once Tool-REX v3 indexing is applied.
Methodology in Plain English
The researchers built a retriever that sits between an agent and a library of skills, and exposed it as an MCP server so different agent harnesses could plug it in without code changes. When the agent asks for a skill, SkillSeek does two things: first, a compact embedding model scores the query against every indexed skill and returns the 20 best matches; second, a slightly larger cross-encoder looks at those 20 query-skill pairs side by side and reranks them, returning the top 5. The agent can then ask for the full body of any skill it wants.
Rather than index each skill's full text, the researchers index only the skill's name and description, plus four short structured tags (file type, primary operation, two secondary operations) generated offline by an LLM. The body text is deliberately left out and fetched only on demand.
They evaluated this against alternatives on two skill pools: a hand-curated 192-skill pool built from the SkillsBench corpus (which pairs 89 verifier-backed tasks with 233 raw SKILL.md files that deduplicate to 192 unique names, roughly 2.6 skills per task, with about 41 skills appearing in more than one task's bundle), and a noisy 34,000-skill marketplace pool harvested from public hubs. Two agent backbones were used — Qwen3.5-397B-A17B as the primary and MiniMax-M2.7 as a weaker secondary, roughly 99 CodeArena points lower — driven through the OpenHands SDK, which was chosen because it is open-source, Python-scriptable, and supports MCP natively. Every trial ran in a fresh Docker container, capped at 30 agent turns and 600 seconds of wall clock, and was graded by each task's deterministic pytest verifier, which returns a graded reward in [0, 1] rather than a binary pass. The reported metric is the mean graded reward over 89 tasks with missing trials counted as 0.
A server-side X-Skill-Method HTTP header let the driver swap retrieval backends between conditions without changing what the agent saw, keeping the comparisons directly comparable. The authors also reimplemented Liu et al. (2026b)'s two variants inside their own driver — liu_hybrid, which runs BM25 fused with Qwen3-Embedding-4B via reciprocal-rank fusion over top-60 inside the main loop, and liu_refined, which moves that exploration into a separate refinement subagent run once per task.
Why This Matters
Impact on research. The paper challenges the standing assumption that agent-skill selection must be delegated to the agent itself. It provides a controlled, cost-instrumented comparison showing the standard IR recipe can be a strong default under the tested conditions, and it contributes the helpfulness-gap diagnostic as a reusable way to separate retrieval recall from downstream agent usefulness — two quantities the authors show can move in opposite directions.
Real-world applications.
- Enterprise agent deployment: teams running MCP-capable coding or workflow agents can add a CPU-only retriever that leaves token spend essentially unchanged from running no skills, instead of paying a per-task LLM exploration tax.
- Skill marketplace platforms: operators of large public skill hubs can offer deterministic retrieval as a low-latency, low-cost selection layer that scales to tens of thousands of skills without per-query LLM calls.
- Security and risk gating: the paper notes adjacent work on pre-load risk scoring (0.800 F1 at six tenths of a US cent per skill) and describes composing it with retrieval, since a deterministic retriever gives a clean insertion point for filtering before a skill reaches the agent.
- Agent harness and SDK development: because the retriever is exposed over MCP, harnesses like Claude Code, Codex CLI, and OpenHands can adopt it without source changes.
Industry relevance. The dominant cost concern for production agents is per-task token spend. A retrieval layer that costs $0 in LLM tokens and adds roughly 1.1 seconds of CPU latency per lookup, while leaving agent-loop spend within fifty cents of the no-skill floor, is directly aligned with that concern. The finding that the agent commits to the rank-1 candidate about 65% of the time also tells reranker vendors and builders that optimizing top-1 usefulness — not top-5 recall — is what moves agent outcomes.
Future Directions
-
Repeated independent rollouts. The headline Qwen3.5 settings were run once per cell; the authors note that establishing equivalence in a stronger sense would require a pre-specified practical margin and repeated rollouts separating task variation from rollout variation.
-
Transfer beyond the tested regime. All results come from the 89 SkillsBench tasks under OpenHands with two backbones and two pools. Terminal, web, and software-engineering agents impose different retrieval demands, and whether the parity holds there is open. The authors name SkillRet — 16,129 public agent skills paired with 4,392 evaluation queries — as a natural target for testing transfer.
-
Comparing first-stage retrievers by agent outcome, not just recall. The recall-ceiling analysis compares stage-1 retrievers by R@5 only. Running the strongest first-stage alternatives through the same cross-encoder and the same agent loop would settle whether higher recall actually translates into higher pass rate.
-
Determining when LLM-mediated refinement genuinely pays. The paper positions LLM-mediated alternatives as reserved for "harder regimes," but the precise boundary — what pool size, pool noise level, or task composition makes agentic synthesis worth $51.30 per trial — is left uncharacterized. Liu et al.'s own documented tensor-parallelism case, where the agent composed a new skill by merging two partially relevant ones, remains the clearest example of the capability a static retriever cannot reproduce.
Target Audience
This paper is written for IR researchers and applied ML engineers working on retrieval-augmented agents, and for practitioners building or operating skill marketplaces and agent harnesses. It will be most useful to readers already comfortable with two-stage retrieval architectures and MCP-based agent tooling, and
Authors’ abstract
Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a $4 \times 11$ grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.