论文

MASBench:面向部分可观测场景的多智能体协作基准

MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability

精选理由

多智能体协作大多在信息全局可见的假设下测,这个基准专门测每个 agent 只能看到局部信息的情况,还开源了代码,做多智能体系统的可以拿来自测。

研究团队发布 MASBench,专门针对部分可观测环境下的 LLM 多智能体协作评测。基准设置 Reasoning、Scheduling、Game 三类渐进式任务,用于评估 Protocol、Memory、Routing 三种协作机制。评测指标包括性能得分、通信成本和成本效益三项确定性指标。论文实验覆盖多种 LLM 后端与机制配置组合,代码已在 GitHub 开源(BUPT-GAMMA/MASBench)。

原文 · arXiv cs.AI

MASBench: Benchmarking LLM-based Multi-Agent Collaboration under Partial Observability

Large language models (LLMs) have progressively evolved into the core of autonomous agents. Building on this progress, LLM-based multi-agent systems (MAS) coordinate multiple agents into a synergistic team to accomplish complex tasks that exceed the capabilities of individual agents. The effectiveness of such systems depends not only on the agents themselves, but also on how collaboration mechanisms are designed and organized. Note that real-world collaboration is typically partially observable, where each agent can only access partial information about the environment due to physical or privacy-related constraints. However, many existing multi-agent benchmarks assume global observability, and leave limited support for systematically evaluating collaboration mechanisms. To bridge this gap, we introduce MASBench, a multi-agent collaboration benchmark designed under partially observable constraints. It is organized into three progressive task categories: Reasoning, Scheduling, and Game. Through this structure, we progressively evaluate three representative collaboration mechanisms: Protocol, Memory, and Routing. MASBench further provides deterministic evaluation metrics, including performance score, communication cost, and cost effectiveness, to characterize both collaboration outcomes and communication overhead. Experiments across diverse LLM backbones and mechanism configurations offer empirical guidance for effective MAS design. Code is available at: https://github.com/BUPT-GAMMA/MASBench