算力才是AI的硬通货。看看K3上线两天就崩了,部署成本百万美金起,未来谁有算力谁说了算。
Deedy指出,Kimi K3上线2天即成为OpenRouter第10大模型,每日处理约1400亿token。但由于基础设施承压,吞吐量从30tok/s降至13tok/s,端到端延迟增至72秒,首token延迟超20秒。部署K3最低需购买8块B300(成本约50万美元),推荐配置GB300 NVL72机架耗资约400万美元。Moonshot及多数美国推理提供商均面临算力瓶颈,Meta拥有约7GW算力但无明确模型规划,成为最大变数。算力需求增长超3个数量级,前沿智能性能按METR任务时间提升32倍,算力锁定者受益显著。
阅读原文: https://t.co/jtAmGyb81g
阅读原文: x.com/deedydas/statu… Deedy @deedydas The story of AI in the next few years is going to be compute: an essay on the future of AI. K3 in 2 days is already #10 on OpenRouter with ~140B tok/day, and it’s infra is crumbling. Throughput is down from 30tok/s to 13tok/s, E2E latency is up to 72s and time to first token is >20s! It would cost a minimum of $500k to buy the 8 B300s it would take to serve even quantized Kimi K3 and ~$4M for the more recommended GB300 NVL72 rack. I don’t think Moonshot has the compute available to scale to their demand! In fact, even the US based inference providers will likely not be able to scale capacity as much as they’d like even if they were to host it: a 2.8T model is no joke. GPU providers (neoclouds etc) are doing 3yr and I recently hear 5yr commits with an ungodly 30% down, and customers are chomping it up. Prices continue to go to the moon. The two big labs, hyperscaler clouds, Grok and Meta have compute deals locked in prior, and the rest are fighting for scraps. Tier 1 neoclouds (coreweave/nebius etc) are rumored to not even small “smaller” customers. Meta is the biggest wildcard here. With ~7GW of compute by eoy 2026 and no clear big model ties, they either get to frontier on their own or can host the most Kimi K3 capacity (unless they sell it to the labs). Even though the price of models has fallen over time, it’s worth noting that the price of frontier has not. 3yrs ago, GPT-4 released at $60/M, o1 at $60/M, Opus 4 at $75/M, GPT5 at $10/M, Fable at $50/M and now Sol at $30/M and K3 at $15/M. Even if you consider K3 frontier, that’s only a 4-5x flux in 3yrs. In that time, frontier demand has increased at least 3+ ooms and frontier intelligence performance has gone 32x at least by task time by METR. Essentially, so long as a) the demand for frontier intelligence continues to grow to near infinity, b) the frontier continues to grow in performance, even as c) if the price of frontier declines a little, the value accrued to frontier grows significantly! And there’s a tremendous bull case for those who have locked up compute if you’re bitter lesson pilled and believe larger models will always be smarter models. 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 0 👀 404 📊 1 ⚡