SkillGym:用 2,756 个技能环境微调后,Qwen3.5-35B-A3B 在 Terminal-Bench 2.1 提升 19.10 分
把技能文档变成 2756 个训练环境直接练进模型权重,微调后的 Qwen3.5-35B-A3B 在终端基准上超过了 Claude Sonnet 4.6,思路很新颖。
一篇新论文提出 SkillGym,把人类撰写的技能文档转化成 2,756 个带代码检查器的训练环境,再用 8,364 条成功轨迹做微调。微调后的 Qwen3.5-35B-A3B 在 Claude Code 中运行时,Terminal-Bench 2.1 提升 19.10 分,SkillsBench v1.1 提升 28.13 分、达到 51.47%,高于 Claude Sonnet 4.6 和 GPT-5.4 Mini 的公开成绩。更有意思的是,不加载技能的微调模型也能赢过把技能放进上下文的基座模型,说明技能可以真正练进权重里。
An interesting idea is to continually train specialized models on skills.
This paper explores that idea.
They propose SkillGym which turns human-written skills into 2,756 training environments with code-based checkers, then fine-tunes on 8,364 successful trajectories.
After fine-tuning, Qwen3.5-35B-A3B running in Claude Code gains 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, where it reaches 51.47%.
That is above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini.
With no skills loaded, the trained model beats the base model that has the skills in context.
Paper: https://t.co/ubb53Qd1Ms
- Gorden Sun12:13原文