Jerry Liu 用“任务复杂度”框架揭示了 AI 初创公司的真正护城河——不是模型本身,而是规范复杂任务的能力。做 AI 产品/创业的团队,看完会重新思考自己的定位。
Jerry Liu 提出用“任务复杂度”(指定任务所需的最小信息位数)来分析 AI 初创公司与前沿实验室的共存空间。低复杂度任务(如总结通话记录)可直接用 Claude 等模型完成;高复杂度任务(如遵循 100 页 SOP 处理生产线偏差)需要软件脚手架和治理框架,初创公司可在此领域发力。任务复杂度与验证难度相关但不相同,例如销售代表代理虽易验证但规范复杂、周期长,保险理赔则既难规范又难验证。随着模型能力提升,任务规范所需信息位数会下降,但知识工作领域仍有巨大机会。
This is a great article on how startups/frontier labs can coexist. Another way to look at this is ta...
This is a great article on how startups/frontier labs can coexist. Another way to look at this is task complexity - the number of bits of information needed to specify a task such that AI can solve the task above a threshold of accuracy: * If the minimum number of bits is low (e.g. summarize call transcript), then you can just prompt Claude Cowork to do it. * If the minimum number of bits is much higher (e.g. follow a 100-page SOP for a production-line deviation) - especially if the task needs to be standardized throughout the org - then the act of specifying the task with the relevant guardrails/auditability/communication becomes much more complex, and it is simply infeasible to expect that an organization can harness the core technology without the software scaffolding in place. Higher complexity task specifications are correlated with how complex it is to verify those tasks, though they aren't necessarily the same. I think both directions are opportunities for AI startups to tackle. ✅ E.g. an e2e sales rep agent is somewhat easy to verify (overattain your number), but task specification of how to actually do it is complex, and the time horizon for running it can take over a year - to see whether the rep can actually hit its number! This means that even if Fable 5 can do it accurately by just giving it a goal, there's lots of opportunities to optimize this workflow to massively reduce cost (in this case it matters for S&M spend) ✅ A lot of tasks are both highly complex to specify and hard to verify e.g. complex insurance claim adjudication. In these cases, the massive bottleneck isn't the model itself, but in the human's ability to even define what good looks like to solve the task at hand. As frontier models get better, the minimum number of bits to specify any task will go down, but IMO there will still be a massive gap for knowledge work that any non-frontier lab company can exploit. sarah guo @saranormous x.com/i/article/2064… 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 1 👀 119 ⚡