Cursor 开源 MoK:NVL72 MoE 训练内核提速 2.37 倍

We're open-sourcing Mixture-of-Kittens (MoK), our MoE training megakernel for NVL72s. It fuses all ...

精选理由

Cursor 把 NVL72 上的 MoE 训练内核开出来了,比公开方案快 2.37 倍,想省训练成本直接用。

AI 摘要

Cursor AI 开源了 Mixture-of-Kittens (MoK) 训练内核。该内核面向 NVIDIA NVL72 平台,将全部 MoE 通信与计算融合为单个确定性内核。相比最强公开基线,MoK 的训练速度最高可提升 2.37 倍。开发者现在可以获得完整代码并用于自己的训练任务。

原文 · Cursor

We're open-sourcing Mixture-of-Kittens (MoK), our MoE training megakernel for NVL72s. It fuses all ...

We're open-sourcing Mixture-of-Kittens (MoK), our MoE training megakernel for NVL72s. It fuses all Mixture-of-Experts communication and computation into a single, fully deterministic kernel, and runs up to 2.37x faster than the strongest public baselines. 💬 73 🔄 104 ❤️ 1436 👀 66379 📊 264 ⚡