Maat:为多智能体 LLM 工作流提供确定性契约治理层
Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows
这篇论文有意思的地方在于它自己推翻了自己的初版结果:审计发现 37% 的暂停是误报,诚实拆穿了基准评分的坑,做多智能体系统的人该看看。
Maat 是一个运行时治理层,在不使用语言模型的情况下,依据版本化工作流契约校验智能体之间的交接。在 6 个受控域工作流(6-15 个智能体、522 次试验)中,治理后运行在五个工作流的评分提升 7.7%-29.1%,模型调用成本在早期可归因暂停场景下降 17-53%。但事后审计发现 7 项基准评分器存在缺陷,人工复查 94 次治理臂暂停中发现 35 次误报(37%),计入后治理臂在 6 个工作流中的 4 个低于未治理臂。结论支持对契约可表达的缺陷做确定性交接校验,同时表明验证器配置和暂停归因本身必须被测试。
Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows
Large-language-model multi-agent systems (LLM-MAS) introduce a characteristic reliability problem: an error produced by one agent can be accepted as context by downstream agents and propagate across the workflow. Many proposed safeguards rely on learned or LLM-based judges whose verdicts are themselves probabilistic; we ask whether a deterministic layer can instead stop contract-detectable handoff defects. We present Maat, a runtime governance layer that validates agent-to-agent handoffs against a versioned workflow contract, or anchor, with no language model in the validation or scoring path. We evaluate it in six controlled domain workflows (6-15 agents, 522 trials) with injected data-level defects and a deterministic seven-check rubric. Version 1 reported gains in all six workflows (2.9-26.5%). A post-publication audit found that three benchmark scorers credited any early halt as a prevented defect. On paired trials where the governed run completed or halted on a finding attributable to a verified defect, the rubric score changes by +7.7% to +29.1% in five workflows and is flat in software development; model-call cost falls 17-53% where attributable halts occur early. A hand review of all 94 governed-arm halts found 35 false alarms (37%), caused by validator defects rather than model behaviour; counting those halts as failed work, the governed arm scores below the ungoverned arm in four of six workflows. The results support deterministic handoff validation for contract-expressible defects and show that validator configuration and halt attribution must themselves be tested; they do not establish universal correctness, hallucination detection, or model-independent effectiveness.