Meta 推出 KernelAgent:多智能体自动优化 Triton GPU 内核
Meta 要在 PyTorch 大会上讲 KernelAgent,用多智能体自动写 Triton 内核,100 个任务上比 torch.compile 快 1.56 倍,写算子的人可以看看。
Meta 的 Kaiming Cheng 将在 PyTorch Conference North America 2026 上介绍 KernelAgent,一个用于编写和优化 GPU 内核的多智能体框架。该框架把 GPU 硬件性能信号引入闭环多智能体工作流,指导 Triton 内核优化。在 KernelBench L1 全部 100 个任务上,KernelAgent 比早期版本生成的内核提速 2.02x,比默认 torch.compile 平均提速 1.56x。
Optimizing kernels can be time-consuming and require deep expert knowledge. KernelAgent is a multi-agent harness that can write and optimize kernels for you.
At #PyTorchCon North America 2026, @KaimingCheng (@Meta) will present “KernelAgent: Hardware-Guided GPU Kernel Optimization via Multi-Agent Orchestration.”
The session will cover a hardware-guided optimization layer that integrates GPU hardware-performance signals into a closed-loop multi-agent workflow to guide optimization for Triton kernels. Across all 100 KernelBench L1 tasks evaluated, KernelAgent achieved a 2.02x speedup over kernels generated by earlier versions and an average 1.56x speedup compared with default torch.compile.
Register for PyTorch Conference North America 2026: https://t.co/jBApW8nESi