这篇论文揭示了自动从交互数据中挖掘技能库的现状:能高精度聚类,但跨域提升有限。看了能知道方法还能改进哪里。
研究提出三阶段流程:分割GUI轨迹、聚类为候选技能、训练技能感知策略。在InteraSkill Workflows基准上,8个聚类中有5个纯度≥0.95。然而,GRPO仅将技能步准确率从18.5%提升到20.5%,BrowseComp+几乎不变,且频率先验在源域指标上更优。表明轨迹挖掘可暴露可检查的技能结构,但当前边界检测器、无序段表示和离线奖励模型不足以可靠跨域策略改进。
Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining
Explicit skill libraries make computer-using agents easier to inspect, but it remains unclear whether such libraries can be mined from interaction data in a way that improves downstream policies. We study this question through a three-stage pipeline that segments GUI trajectories, clusters segments into candidate skills, and trains a skill-aware policy from the resulting annotations. The mined clusters are readable on the source benchmark: five of eight clusters have at least 0.95 purity against InteraSkill Workflows labels. However, readability does not imply transfer. GRPO improves IW skill-step accuracy only from 18.5\% to 20.5\%, leaves BrowseComp+ essentially unchanged, and underperforms trivial frequency priors on key source-domain metrics. We therefore present the method as a diagnostic study: trajectory mining can expose inspectable skill structure, but the current boundary detector, orderless segment representation, and offline reward model are insufficient for reliable cross-domain policy improvement.