KernelZero:Proposer 与 Coder 协同进化生成高性能 GPU Kernel
KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation
一个 7B 模型在 CUDA 上干过 Claude-4.5-Sonnet,靠的是两个小模型互相出题互相训练,生成 GPU kernel 的可以看看这套思路。
KernelZero 是一个 GPU kernel 自动生成框架,用两个模型分工协作:Proposer 从 API 集合生成 Torch 模块,Coder 将其翻译为 CUDA 或 Triton kernel。框架会根据 Coder 当前的薄弱点持续生成与其能力匹配的训练模块,并提出 CA-GRPO 算法,先保证正确性再优化性能。KernelZero-7B 在 CUDA 上超过 Claude-4.5-Sonnet,在 Triton 上超过 DeepSeek-V4-Pro。在 KernelBench Level 1 和 Level 2 上,CUDA pass@1 分别达到 75.8% 和 69.6%,Triton pass@1 分别达到 77.2% 和 72.5%。
KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation
High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality training data aligned with the model's current capabilities, and the inherent trade-off between kernel correctness and performance. To address these challenges, we propose KernelZero, a co-evolution framework that continuously improves GPU kernel generation through two specialized models: a Proposer that generates Torch modules from API sets and a Coder that translates them into CUDA or Triton kernels. KernelZero uses a frontier-driven module generation mechanism to continuously produce capability-aligned training modules based on the Coder's current weaknesses. It further introduces Correctness-Aware Group Relative Policy Optimization (CA-GRPO), which optimizes performance only after correctness becomes sufficiently reliable. By alternating the optimization of the Proposer and Coder, KernelZero forms an automatic curriculum that enables targeted and training-efficient capability improvement. Empirically, KernelZero-7B surpasses Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton. On KernelBench Level 1 and 2, it achieves CUDA pass@1 scores of 75.8% and 69.6%, respectively, with pass@10 reaching 100% and 97%. On Triton, it achieves pass@1 scores of 77.2% and 72.5%, respectively.