这篇论文分析了智能体AI在不确定环境中的失败机制,还给出了SCI和DMM实用框架。如果你做AI智能体开发,这些形式化结论值得参考。
Agentic AI任务在长链执行时因环境不确定性呈指数级失败,每步确定性δ<1时k步成功率衰减为δ^k。论文提出三个形式化结果:确定性-效率界限、验证者-古德哈特定理下限、环境技能演化的收敛条件。研究者构建了基于五个可测量属性的供应确定性指数(SCI)和五级确定性成熟度模型(DMM)。论文还提出了一个可证伪的开放问题框架OQ1-OQ5。立场与平台无关,并讨论了模拟到现实充分性、对齐充分性和AI作为正常技术三种竞争观点。
Grounded Scaling: Why Agentic AI Needs Deterministic Environments
Long-chain agent execution fails exponentially in environments designed for human tolerance: with per-step determinism $δ< 1$, $k$-step chain success degrades as $δ^k$. The AGI-to-ASI scaling debate (Genewein et al., 2026) has so far framed progress as a race between compute growth and a list of frictions (data wall, abstraction barrier, embodied bottleneck, multi-agent trust); we argue that environment determinism is a complementary binding axis cutting across all four, for the broad class of agentic AI tasks whose outcomes are verifiable economically, physically, or through multi-party settlement. Three formal results pin down the regime: a Determinism-Efficiency Bound on chain-task success, a Verifier-Goodharting Floor on flywheel ceilings under imperfect rewards, and a convergence condition for environment-side skill evolution. We operationalise the framework as a Supply Certainty Index (SCI) over five measurable properties, a five-level Determinism Maturity Model (DMM) as adoption ladder, and a falsifiable open-question programme (OQ1-OQ5) with explicit null results that would force retraction. The position is platform-agnostic. We engage three competing positions: sim-to-real sufficiency, alignment sufficiency, and AI-as-normal-technology.