想看芯片和AI效率的前沿研究?这场YC研讨会涵盖了多GPU内核优化、智能每瓦特指标和异构推理硬件,都是干货。
YC Paper Club 最新研讨会聚焦多GPU内核编程、每瓦特智能效率及异构推理。Francois Chabaud 探讨芯片与内核专化;Stuart Sutherland 介绍 Parallel Kittens 项目,简化多GPU AI 内核的设计。Jon Saad-Falcon 提出“智能每瓦特”指标,评测本地与云端AI模型效率。Mark Saroufim 展示AI编写系统代码的进展,Misha Smelyanskiy 强调异构硬件对推理的重要性。
At our latest YC Paper Club, researchers and builders presented on multi-GPU kernels, intelligence p...
At our latest YC Paper Club, researchers and builders presented on multi-GPU kernels, intelligence per watt, heterogeneous inference, and more. Thank you to our presenters: 0:00 – @FrancoisChauba1 : The case for chip and kernel specialization 7:16 – @stuart_sul : Parallel Kittens - Systematic and Practical Simplification of Multi-GPU Al Kernels ( arxiv.org/abs/2511.13940 ) 21:29 – @JonSaadFalcon : Intelligence per Watt - Measuring the Intelligence Efficiency of Local and Cloud AI ( arxiv.org/abs/2511.07885 ) 31:05 – @MarkSaroufim : When Al Starts Writing Systems Code 47:04 – Misha Smelyanskiy: Why AI Inference Needs Heterogeneous Hardware 1:04:33 – @shacklettbp : Building a High-Throughput Game Engine that Runs ENTIRELY on the GPU ( madrona-engine.github.io/shacklett_sigg… ) Your browser does not support the video tag. 🔗 View on Twitter 💬 5 🔄 7 ❤️ 46 👀 6993 📊 12 ⚡