AI模型精选72°

Cursor开源MoK:MoE训练内核速度提升2.37倍

dont worry kitten

精选理由

Cursor开源了MoK内核,把MoE训练的通信和计算合成一个内核,跑得比公开基线快2.37倍,搞NVL72训练的可以看看。

AI 摘要

Cursor宣布开源Mixture-of-Kittens (MoK),这是一个面向NVL72的MoE训练megakernel。MoK将专家混合的所有通信与计算融合进单个确定性内核,相比最强公开基线提速最高2.37倍。该内核专为NVL72集群设计,用于提升大规模MoE训练效率。

原文 · eric zakariasson

dont worry kitten

dont worry kitten Cursor @cursor_ai We're open-sourcing Mixture-of-Kittens (MoK), our MoE training megakernel for NVL72s. It fuses all Mixture-of-Experts communication and computation into a single, fully deterministic kernel, and runs up to 2.37x faster than the strongest public baselines. 🔗 View Quoted Tweet 💬 4 🔄 1 ❤️ 55 👀 4263 📊 8 ⚡