OpenAI、Google和Nvidia都在竞相提升AI模型推理速度,各有不同技术方案。
OpenAI展示GPT 5.6 Sol每秒处理750个token。Google发布Gemini 3.7 Flash平均每秒330个token。Nvidia推出Nemotron 3.5 Lightning并配备NeMo Switchyard动态步进路由技术。更快的吞吐量和更低的延迟可减轻开发者上下文切换负担。这些模型支持实时智能体工作流。
⚡ Top AI companies think inference speed is an architectural requirement worth paying for. OpenAI a...
⚡ Top AI companies think inference speed is an architectural requirement worth paying for. OpenAI and Cerebras demonstrated GPT 5.6 Sol running at 750 tokens per second. Google released Gemini 3.7 Flash averaging 330 tokens per second. Nvidia launched Nemotron 3.5 Lightning with NeMo Switchyard for dynamic step routing. Faster throughput and lower latency alleviate developer context switching and power real-time agentic workflows. Read the complete breakdown in The Batch: hubs.la/Q04w6R0y0 📖 #DeepLearningAI I #AI I #TechNews s 💬 4 🔄 0 ❤️ 9 👀 1557 📊 4 ⚡