Cursor把训练内核MoK开源了,在GB300 NVL72上比最强公开基线快2.37倍,但得先有NVL72机器才能跑。
Cursor Research开源了MoK,这是支撑Composer模型的MoE训练内核。MoK将专家混合的全部通信与计算融合进单个确定性内核。在GB300 NVL72机架上,MoK比最强公开基线快2.37倍。它需要Blackwell SM100或SM103 GPU,没有NVL72算力的用户无法使用。
Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks
Cursor Research has open-sourced Mixture-of-Kittens (MoK), the MoE training megakernel behind its Composer models. MoK fuses all mixture-of-experts communication and computation into a single deterministic kernel, and runs up to 2.37x faster than the strongest public baseline on GB300 NVL72 racks. It requires Blackwell SM100 or SM103 GPUs, which puts it out of reach for anyone without NVL72 capacity. The post Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks appeared first on MarkTechPost .
- Cursor16:00原文