高盛研究:token 需求增长18倍,但 GPU 需求未必同步
高盛这份数据挺反直觉:token 涨18倍,GPU 用量可能远低于预期,因为增量都流向了小模型和开源模型。做基建判断的值得看看。
高盛研究预计总 token 需求从 2025 年 12 月到 2026 年 9 月增长 18 倍,而前沿模型 token 需求仅增长约 8-9 倍。增量 token 越来越多来自更便宜的小模型、开源模型或路由模型,而非计算量最大的前沿模型。量化、蒸馏、投机解码、缓存优化等技术进一步降低了每个 token 消耗的 GPU 算力。因此判断 AI 基础设施周期不能只看 token 增长,还要看 token 量、token 价格与每 token 算力需求的共同作用。
AI is entering a volume economy: because token demand now has to outrun the price collapse underneath it.
Per Goldman Sachs research, total token demand is accelerating (by 18x from December-2025 to Sept-2026 ), while frontier-model token demand is growing much more slowly, at roughly 8-9x on the same index.
the bar for sustaining investment spending getting higher as inference becomes commoditized.
That means the marginal AI token is increasingly coming from cheaper, smaller, open-source, or routed models, not necessarily from the most compute-heavy frontier model. This changes the infrastructure math quite a bit.
That gap is so important for hyperscaler's capex-cycle, because frontier models remain an important source of hyperscaler compute demand, while increasingly competitive open-source models are helping push average token prices down.
So token growth alone is no longer enough to read the AI infrastructure cycle. The relevant variable is the interaction between token volume, token pricing, and the amount of compute required to serve each token.
What really matters is:
tokens consumed × compute required per token
If token volume grows 18x but more of those tokens move to models that need far fewer GPU cycles per token, total GPU demand can grow much slower than token demand suggests. Quantization, distillation, speculative decoding, better caching, smaller models, and model routing all push in that direction.
- PolymarketMoney09-25 12:52原文