ThinkingCap微调Qwen 27B实现2倍推理Token减少

i like this idea of fine-tuning LLMs for efficient reasoning, especially when the intervention remai...

精选理由

想给模型推理提速?这个ThinkingCap把Qwen 27B的思考token砍掉一半,有的例子快10倍,效果还很稳定。

AI 摘要

Jaroslav Beck团队发布ThinkingCap高效模型系列,在Qwen 3.6 27B模型上进行微调,平均减少2倍思考token,部分示例生成速度提升10倍。该方法保持非侵入性干预,得到的模型与原检查点行为高度相似。作者认为这个技术可能像量化一样成为默认工具包的一部分。

原文 · Thomas Wolf

i like this idea of fine-tuning LLMs for efficient reasoning, especially when the intervention remai...

i like this idea of fine-tuning LLMs for efficient reasoning, especially when the intervention remains as non-invasive as possible and the resulting model behaves very similarly to the original checkpoint wondering if it could become part of the default toolbox in the field, like quantization as become Pic: in (left) and out (right) of domain behavior Jaroslav Beck @JaroslavBeck Today we’re announcing our “ThinkingCap” efficient model series with a 2× thinking token reduction on average in Qwen 3.6 27B, with up to 10x faster generation on individual examples. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 6 🔄 1 ❤️ 5 👀 929 📊 5 ⚡