GPT-5.6 自己优化自己,服务成本降了20%,生成效率提了15%,省下来的钱都能买好几个数据中心了。
GPT-5.6 通过 Codex 分析生产流量、改进负载均衡、重写 GPU kernel 和运行数百次推测解码实验,将端到端服务成本降低 20%。推测解码使 token 生成效率提升超 15%。这些优化为 OpenAI 每月节省数十亿美元成本。
GPT-5.6 found optimizations that "reduced end-to-end serving costs by 20%" for OpenAI to serve that ...
GPT-5.6 found optimizations that "reduced end-to-end serving costs by 20%" for OpenAI to serve that model Presumably that's billions of dollars a month in savings at this point? Vaibhav (VB) Srivastav @reach_vb Codex analysed production traffic, improved load balancing, rewrote production GPU kernels and ran hundreds of experiments on its own speculative-decoding model. The kernel improvements reduced end-to-end serving costs by 20%, while speculative decoding improved token-generation efficiency by more than 15%. 🔗 View Quoted Tweet 💬 7 🔄 3 ❤️ 32 👀 5322 📊 8 ⚡