DeepMind真会玩:把模型权重当文本喂给LLM,不用训练就能直接生成新技能,论文在arXiv可看。
谷歌DeepMind在arXiv:2607.27497提出SkillSmith,把模型权重当作额外模态让LLM原生读取。SkillSmith增强后的模型摄入现有前缀权重和描述能力的文本,直接输出体现该技能的新前缀权重。SkillSmith让技能组合变成推理期操作,不再需要训练。论文指出,SkillSmith的效果超过单独文本或单独权重适配。
New research from Google DeepMind. (bookmark it) SkillSmith treats model weights as an additional ...
New research from Google DeepMind. (bookmark it) SkillSmith treats model weights as an additional modality the LLM reads natively. The augmented model ingests existing prefix weights alongside rich text describing how a capability relates to a target, then directly outputs new prefix weights that manifest that skill. Skill composition becomes an inference-time operation instead of a training run. The team calls this instruction-steered parametric synthesis. The gains exceed what text-only and weight-only adaptation reach on their own. Paper: arxiv.org/abs/2607.27497 Track more trending AI papers in our academy: academy.dair.ai 💬 3 🔄 3 ❤️ 23 👀 2236 📊 7 ⚡