AI模型精选

GPT-5.6 优化降低端到端服务成本20%,效率提升15%

GPT-5.6 found optimizations that "reduced end-to-end serving costs by 20%" for OpenAI to serve that ...

精选理由

GPT-5.6 自己优化自己,服务成本降了20%,生成效率提了15%,省下来的钱都能买好几个数据中心了。

AI 摘要

GPT-5.6 通过 Codex 分析生产流量、改进负载均衡、重写 GPU kernel 和运行数百次推测解码实验,将端到端服务成本降低 20%。推测解码使 token 生成效率提升超 15%。这些优化为 OpenAI 每月节省数十亿美元成本。

原文 · Simon Willison

GPT-5.6 found optimizations that "reduced end-to-end serving costs by 20%" for OpenAI to serve that ...

GPT-5.6 found optimizations that "reduced end-to-end serving costs by 20%" for OpenAI to serve that model Presumably that's billions of dollars a month in savings at this point? Vaibhav (VB) Srivastav @reach_vb Codex analysed production traffic, improved load balancing, rewrote production GPU kernels and ran hundreds of experiments on its own speculative-decoding model. The kernel improvements reduced end-to-end serving costs by 20%, while speculative decoding improved token-generation efficiency by more than 15%. 🔗 View Quoted Tweet 💬 7 🔄 3 ❤️ 32 👀 5322 📊 8 ⚡

GPT-5.6 优化降低端到端服务成本20%,效率提升15% · AI 热点