Fireworks 和 MiniMax 把稀疏注意力内核开源了,实测吞吐快了1.6倍,搞推理优化的不可错过。
Fireworks AI 和 MiniMax 联合公开了 MiniMax Sparse Attention(MSA)内核的代码仓库。Fireworks 优化了注意力内核的加载与存储流水线,实现了 1.6 倍吞吐量提升。两个内核仓库已分别在 GitHub 上开源(fw-ai/minimax-… 和 MiniMax-AI/MSA…)。这项工作旨在推动开源实现的实际价值。
Weights aren't the only thing that should be open-source. Both kernel repos behind our recent work ...
Weights aren't the only thing that should be open-source. Both kernel repos behind our recent work with @MiniMax_AI are now public: → Fireworks: github.com/fw-ai/minimax-… → Minimax: github.com/MiniMax-AI/MSA… RyanLee @RyanLeeMiniMax Great to see @FireworksAI_HQ continuous optimizations on MiniMax Sparse Attention (MSA). By refining the attention kernel’s load and store pipelines, they’ve achieved a 1.6x throughput uplift. This work should bring meaningful value to open-source implementations. fireworks.ai/blog/kernel-op… 🔗 View Quoted Tweet 💬 0 🔄 1 ❤️ 20 👀 1712 📊 3 ⚡