AMAO 框架用多智能体处理运筹学问题中的逻辑不一致
Ambiguity-Aware Multi-Agent Framework for Automated Operations Research under Logical Inconsistency
一篇做自动运筹学建模的论文,用多智能体先修补含糊描述再建模,8B 小模型跑赢了 Claude-sonnet-4.5。
论文提出 AMAO 框架,针对运筹学问题描述含糊、逻辑不一致的情况,先用逻辑对齐智能体做两阶段对齐,再由源充分性验证器判断证据是否足够,不足时由交互修复智能体向用户追问。32B 模型达到 86.8% 错误恢复成功率和 60.4% 求解准确率,是基座的 1.98 倍和 1.86 倍。8B 模型取得 47.9% 求解准确率,超过 DeepSeek-V4-Flash 和 Claude-sonnet-4.5 等受评变体。在需要澄清的案例上,交互修复将完整描述准确率提升到 90.71%。
Ambiguity-Aware Multi-Agent Framework for Automated Operations Research under Logical Inconsistency
Operations research (OR) problems are often described by stakeholders with incomplete knowledge and vague expressions in real-world settings. Such ambiguous descriptions can not be used to formulation directly. Thus, automating OR problems solving requires processing this logical inconsistency in advance. To address this challenge, we propose an \textbf{A}mbiguity-Aware \textbf{M}ulti-Agent Framework for \textbf{A}utomated \textbf{O}R Problem Solving, \textbf{AMAO}, which first addresses logical inconsistency through two-stage alignment before downstream formulation and coding. Specifically, a logical alignment agent with OR-guided experts structure first produces an aligned candidate using supervised routing over variables, parameters, objectives, and constraints. Then, a source-sufficiency verifier determines whether the original description supports the required repairs. When evidence is insufficient, an interactive repair agent requests additional information and revises the description before modeling and coding. In addition, an ambiguity-aware benchmark and its matched dialogue extension are proposed to support evaluation of both stages. The 32B model achieved 86.8\% error-recovery success and 60.4\% solution accuracy, reaching 1.98 and 1.86 times the respective base-model scores. The 8B model achieved 47.9\% solution accuracy, exceeding the best evaluated variant such as DeepSeek-V4-Flash and Claude-sonnet-4.5. Interactive repair further achieved 90.71\% complete-description accuracy when cases requiring clarification. These results support combining context-based repair with evidence assessment and targeted user clarification for automated OR modeling.