HySTAR:用锚定超图改进多智能体强化学习的信用分配
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
做 MARL 的可以看看这篇,用固定超图解决信用分配漂移,SMAC 上比 MAPPO 提了 16.7%,实验覆盖挺全。
HySTAR 是一个基于 MAPPO 的合作式多智能体强化学习框架,针对共享奖励下团队收益难以分配到个体的结构化目标漂移问题,用固定锚定的稀疏超图作为时间一致的值分解基础。方法结合时空编码器表示交互,并按时间与结构相关性构造各智能体专属优势。在 SMAC 最难设定上比 MAPPO 相对提升 16.7%,比 HYGMA 提升 15.6%,在全部六个 GRF 场景排名第一,Traffic Junction 相比 MAGIC 收敛轮数最多减少 40.2%。
HySTAR: Anchored Hypergraphs for Stable Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
Cooperative multi-agent reinforcement learning under partial observability and shared rewards requires assigning team outcomes to individual agents and high-order coalitions. A MAPPO-style critic compresses joint behavior into one global value, while critics that dynamically reconstruct the grouping topology change the mapping from agents and coalitions to value components as interactions or active agents evolve. We refer to this inconsistency as structural target drift. We introduce HySTAR, a MAPPO-based framework that separates adaptive representation learning from a temporally consistent high-order value-decomposition basis. HySTAR anchors an overlapping sparse hypergraph as a uniformly covered decomposition scaffold, uses a spatiotemporal encoder to represent physical and task-dependent interactions, and combines temporal and structural relevance to construct agent-specific advantages. Experiments on SMAC, GRF, Traffic Junction, and MPE demonstrate consistent improvements over MAPPO-style, value-factorization, and dynamic-grouping baselines. On the hardest SMAC settings, HySTAR achieves relative gains of 16.7\% over MAPPO and 15.6\% over HYGMA, ranks first on all six GRF scenarios, reduces Traffic Junction convergence epochs by up to 40.2\% relative to MAGIC, and obtains the highest MPE episode rewards. Controlled topology, agent-death, neighborhood, and parameter analyses support the benefit of anchoring the decomposition scaffold while adapting the propagated representations.