Papers
arxiv:2609.38822

SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

Published on Sep 30
ยท Submitted by
David Yang
on Oct 1
Authors:
,
,
,

Abstract

Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a 4 times 11 grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.

Community

Paper submitter

Accepted at AACL-IJCNLP 2026, Findings.

An Agent Skill is a folder with instructions that teach an AI agent one specific job. The agent reads a short summary of each skill at startup and opens the full text only when the job looks relevant.

Public collections of these skills have grown fast. The most recent crawl counts 238,180 across the major distribution platforms.

So anyone can write a skill. Choosing which ones to load is the hard part.

โŒ Loading everything does not work. At a pool of only 192 skills, loading all of them leaves the agent exactly at its no-skill score while raising token use by 59%.

โŒ Loading loosely does not work either. One relevant skill lifts the agent by 17.8 points. Two or three lift it by 18.6. Four or more drop it back to 5.9.

So a retriever has to be right, not merely reasonable.

The answer the field has settled on is to let the agent retrieve for itself: rewrite the query, look through candidates, and combine them into a new skill for the task.

That answer has a real strength. On one task about splitting tensors across GPUs, the agent found two partly relevant skills and merged them into something neither one contained. A fixed retriever cannot do that.

๐Ÿ” So we asked how often that strength decides the outcome.

We ran 11 retrieval methods across 4 settings, two pool sizes crossed with two language models, on an 89-task benchmark where a test suite decides whether the agent succeeded.

๐Ÿ“Š Results. Pass rate is the share of the 89 tasks solved, so higher is better.

โœ… BM25 alone, a keyword-matching method from the 1990s, matches or beats the agentic loop on 3 of the 4 settings
โœ… A small cross-encoder closes the gap on the fourth, reaching the same 0.442 pass rate
โœ… Cost per run: USD 27.41 with no skills, USD 27.54 with our retriever, USD 51.30 with the agentic loop
โœ… Our retriever runs on CPU and spends zero language model tokens

We are not arguing that agentic retrieval is wrong. We are arguing it should be the fallback, and the standard retrieval recipe should be what you try first.

โš ๏ธ One limit we report in the paper: on the large pool, first-stage recall caps what any reranker can do, and that ceiling is where the agentic loop earns its cost.

๐Ÿ™ Joint work with Wenlong Zhang, Tian Shi, and Ping Wang.

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.38822 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.38822 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.38822 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.