模型路由是 AI 应用降本增效的关键,做 AI 产品、智能体或基础设施的团队值得关注——它可能成为下一个像 API 网关一样的基础设施层。
Jerry Liu(LlamaIndex 创始人)认为,AI 创业公司将在“模型路由即服务”领域积累大量价值,这不仅是 OpenRouter 这样的通用路由,还包括垂直化的智能体和基础设施。他以文档基础设施(解析、提取、搜索)和网络搜索(Exa/Parallel)为例,说明在准确性与成本的帕累托曲线上找到最佳点既重要又困难。Brian Armstrong 补充说,未来 80% 的工作负载将运行在便宜 99% 的模型上,只有 20% 需要最新高端模型,而 Coinbase 已通过路由提示词到更便宜的模型来保持成本稳定。这揭示了模型路由作为降低 AI 应用成本、提升效率的关键基础设施,对开发者和创业公司是巨大机会。
I really do think we'll see a lot of value accrue in AI startups building "model routing as a servic...
I really do think we'll see a lot of value accrue in AI startups building "model routing as a service" Not just OpenRouter - this includes a much broader set of verticalized agents and infrastructure. * For us it's document infrastructure: parsing, extraction, search * Another infra analogy is web search: Exa/Parallel * For verticalized apps it could be anything from Cognition to Harvey The frontier labs own the underlying models, and their main application at the app layer is the enormous amounts of $$ that they have; but their main disadvantage is that they only own a subset of all the points on the Pareto curve. Speaking from our team's experience, it is both non-trivial and extremely important to find a point on the pareto curve of accuracy and cost. There's enormous amounts of ML time spent on model evaluations and benchmarking, and infra time making sure that the service can scale reliability, without rate limits, and without blowing up cost. At the same time it's extremely important not just for cost reasons but also latency, and oftentimes long-tail accuracy in use cases that demand it. Brian Armstrong @brian_armstrong Good take My guess is - demand for intelligence is near infinite - but 80% of workloads will be running on 99% cheaper models within 12-18 months - 20% of workloads will still run on latest gen models where IQ maxing is important (scientific breakthroughs, higher level ochestrator agents?) - rough analogy might be what % of macbooks or gaming PCs sold have the maxed out specs for CPU/GPU, prices are falling much faster than Moore's law here though - this leads me to think the limiting factor will be energy and compute, not better models At Coinbase we're working hard on routing prompts to cheaper models where appropriate, and in some cases have been able to keep costs roughly flat, while token usage continues to grow exponentially. 🔗 View Quoted Tweet 💬 5 🔄 0 ❤️ 4 👀 375 📊 5 ⚡