Jerry Liu 用真实成本数据揭示了模型选择的巨大经济差异,做 AI 应用选型或成本控制的团队值得仔细看——选对模型能省下 20-40 倍 token 成本。
LlamaIndex 创始人 Jerry Liu 指出,没有前沿实验室能独占成本、延迟与精度的帕累托前沿所有点,开源模型在成本上可低数个数量级。他观察到组织对模型路由和成本优化的兴趣激增,原因包括企业更谨慎管理成本,以及 AI 初创公司寻求构建护城河和提高毛利率。他引用 Chamath 的数据对比:每月 10 亿 token 输入/输出场景下,GPT-5.5 Pro 成本约 10.5 万美元,而 DeepSeek V4 Pro 仅需 5220 美元,能力差距远小于价格差距。Jerry 认为,随着控制平面(如 Software Factory)普及,前沿实验室收入增速将下降,开源模型收入将飙升。
No frontier lab will own every single point on the pareto frontier around cost/latency and accuracy....
No frontier lab will own every single point on the pareto frontier around cost/latency and accuracy. Even as the pareto frontier itself advances, there will always be points owned by open-weight models that are orders of magnitude cheaper than the frontier ones. There's been this huge uptick in interest in model routing and cost optimization, for two reasons: ✅ Organizations are more carefully thinking about how to carefully manage cost ✅ Every AI-native startup and VC is thinking about the best way to build a moat against the frontier labs (and raise their gross margins) These topics are quite relevant for our mission at @llama_index - which is to build the document infrastructure for AI agents. We want to help unlock the trillions of pages of unstructured paperwork within every organization for agentic automation. Organizations require both higher accuracy and orders of magnitude lower cost within their document OCR solutions compared to the frontier models. The pareto frontier of document OCR has a meaningful, exploitable gap beyond what is offered by the frontier VLMs. Exploiting this gap requires both AI expertise and an absolutely obsessive focus around how PDFs, .docx, .pptx, and other file formats work. Chamath Palihapitiya @chamath Your margin is my opportunity: AI version… The biggest surprise of 2026 is that the capability gap between the best open-weight/source models and the best closed models has narrowed much faster than the pricing gap. The pricing gap remains enormous while the capability gap is quite narrow. What does this means in practice? For a company consuming 1 billion input tokens and 1 billion output tokens per month: GPT-5.5 Pro: ~$105,000 Claude Opus 4.8: ~$30,000 DeepSeek V4 Pro: ~$5,220 DeepSeek R1: ~$2,740 I asked ChatGPT what it thought about this and it answered as follows: “If I were building a company today, the economic frontier would look roughly like: DeepSeek V4 Pro / R1 for high-volume inference. Claude Opus for premium agent workflows where reliability matters. GPT-5.5 Pro only for workloads where its incremental capability demonstrably produces enough business value to justify a 20–40× token premium.” Most CEOs have no idea that, instead of this nuanced approach, their teams are running amok internally by picking the most expensive models in most cases and burning through massive budgets with zero governance, audit ability and control. As control planes like our Software Factory become more standard, you can expect the run rate revenue growth of the frontier labs to go down meaningfully and the revenues of the open models to skyrocket. Why? Because we can implement the nuanced approach above and be agnostic to model - instead focusing on customer intent, model task and cost management among other things. 🔗 View Quoted Tweet 💬 7 🔄 1 ❤️ 8 👀 756 📊 7 ⚡