斯坦福大学AI实验室提出Prefix Sliding,优化长任务中全注意力机制的效率,值得一读。
Muennighoff等提出Prefix Sliding方法,通过遗忘部分上下文信息,提高长任务中全注意力机制的效率,避免信息丢失。新论文发表于斯坦福大学AI实验室。
Long reasoning traces are expensive because full attention grows with context length. @Muennighoff ...
Long reasoning traces are expensive because full attention grows with context length. @Muennighoff et al find that maybe LLMs can stand to forget a little… with Prefix Sliding! Niklas Muennighoff @Muennighoff new paper: Prefix Sliding for efficient test-time scaling vanilla full attention OOMs on long tasks & compaction loses important details -- prefix sliding is a simple & fast alternative that can outperform both 📜 https://t.co/fUw7yJAN5D 🔗 View Quoted Tweet 💬 2 🔄 0 ❤️ 13 👀 1468 📊 3 ⚡
- arXiv cs.AI08-26 17:37原文