论文

去中心化团队决策新方法:低秩POMDP下无需中心化训练

A Decentralized Partially Observable Team Decision Methodology with Delayed Information Sharing

精选理由

多智能体协作不用中心化训练也能逼近团队最优解,还带样本复杂度证明,做分布式决策的可以看看。

论文研究带低秩潜在动态的去中心化部分可观测团队决策问题,不假设已知转移模型。每个成员基于本地私有信息和延迟共享的公共信息,学习一个近似低秩MDP,并用最小二乘值迭代计算策略。该方法无需中心化协调者或中心化训练,仍能逼近中心化团队最优解的对应分量。论文还给出了有限样本性能保证和样本复杂度界。

原文 · arXiv cs.LG

A Decentralized Partially Observable Team Decision Methodology with Delayed Information Sharing

We study decentralized partially observable team decision problems with low-rank latent dynamics and unknown system models. The proposed framework combines team-theoretic equivalence with low-rank model representations to address cooperative decision-making in partially observable Markov decision processes without prior knowledge of the transition model. Each team member makes decisions based on local private information and delayed common information shared across the team. Using only this available information, each member learns an approximate low-rank Markov decision process and applies least-squares value iteration to compute its policy. This yields a fully decentralized learning and planning algorithm that requires neither a centralized coordinator nor centralized training. We show that the resulting member-side solutions approximate the centralized team solution: despite partial observability, unknown dynamics, and delayed common information, each member recovers the corresponding component of an approximate team-optimal policy. We further establish finite-sample performance guarantees and derive a corresponding sample-complexity bound for the proposed algorithm.