行业71°

vLLM任意时刻都在50万张GPU上运行,却少有人知

vLLM runs on half a million GPUs at any given moment. Most people have never heard of it. Simon Mo,...

精选理由

vLLM同时占着50万张GPU,却是开源模型的隐身基座。听Simon Mo拆解Kimi K3和开源模型上生产的坑。

AI 摘要

vLLM是一个开源推理引擎,任意时刻运行在约50万张GPU上,构成开源模型生产环境的隐形基础设施。Simon Mo作为vLLM维护者和Inferact CEO,在a16z播客中讨论了开放模型的优势、发布首日模型适配的幕后冲突,以及Kimi K3实际带来的提升。他还对比了OpenRouter、Ollama等工具在ChatGPT出现前的起源,并解释了开源模型许可证正在变化的原因。访谈也涉及如何基于开源项目建立公司。

图片来源 · a16z
原文 · a16z

vLLM runs on half a million GPUs at any given moment. Most people have never heard of it. Simon Mo,...

vLLM runs on half a million GPUs at any given moment. Most people have never heard of it. Simon Mo, co-founder and CEO of @inferact and lead maintainer of vLLM, sits down with a16z’s Matt Bornstein and Elena Burger to discuss what it takes to actually run open models in production, the advantages of open models, Simon’s mission at Inferact, and more. 00:00 Intro 01:46 When open source became critical infrastructure 08:55 Day zero model releases, and the drama behind them 14:59 What Kimi K3 actually buys you 18:56 Why open model licenses are changing 22:24 The pharmaceutical analogy for funding model training 26:16 If GPUs got 99% cheaper 29:08 Why vLLM, OpenRouter, and Ollama all started before ChatGPT 35:42 Building a company on an open source project 39:48 The inventor of RoPE removing RoPE @simon_mo_ @BornsteinMatt @VirtualElena Your browser does not support the video tag. 🔗 View on Twitter 💬 4 🔄 3 ❤️ 19 👀 4622 📊 6 ⚡