技巧多源确认

SkillSeek:智能体技能检索新方案

SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

精选理由

Anthropic开源SkillSeek,用传统IR方法替代LLM循环,大幅降低技能检索成本,性能相当。

Anthropic的Agent Skills包已包含23万项可重用技能,选择成为瓶颈。SkillSeek是一个开源的两阶段技能检索器,基于标准IR配方构建。在89项任务的SkillsBench基准测试中,SkillSeek在4×11种配置组合下达到了与Liu等人LLM中介循环相当的性能。每试验成本从51.30美元降至27.54美元,接近无技能基线。

原文 · arXiv: Anthropic

SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale

Anthropic's Agent Skills package reusable procedural know-how for an LLM agent into SKILL.md directories, and open-source aggregations have grown past 230,000 skills, making selection rather than authoring the bottleneck. The standing answer in the literature outsources selection to the agent itself: an LLM-mediated retrieval loop that rewrites queries and refines candidates inside the agent's decision loop, paying LLM tokens on every task. We present SkillSeek, an open-source two-stage skill retriever built from the standard IR recipe (a BGE-base bi-encoder feeding a small cross-encoder, exposed over MCP). Across a $4 \times 11$ grid of pool, backbone, and method on the 89-task SkillsBench benchmark, SkillSeek reaches observed parity with the LLM-mediated loop of Liu et al. at essentially no extra cost: plain bm25 alone records a pass rate at or above their refined loop on three of four settings, and a small cross-encoder covers the remaining difference on the fourth. A first-stage recall ceiling explains the pattern, and total per-trial spend drops from USD 51.30 to USD 27.54 (within fifty cents of the no-skill baseline). Under the SkillsBench tasks and OpenHands harness we tested, this positions the standard IR recipe as a strong default for agent-skill retrieval, with LLM-mediated alternatives a natural fit for cases where deterministic methods fall short.