论文

G2MAF:测试时梯度引导优化多智能体流策略

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

精选理由

多智能体离线RL的新做法:部署后不用重训,靠critic梯度在测试时微调动作,MPE和SMAC上平均提升约9%,延迟只多6%。

论文提出 Gradient Guided Multi Agent Flow(G2MAF),用于优化离线多智能体强化学习部署时冻结的策略。它在测试时用全局归一化的投影 critic 梯度协调所有智能体的动作修正,保持动作可行且接近原策略。在 24 个 MPE 和 SMAC 设置中,标准版本改进了 20 个冻结策略场景,MPE 平均相对提升 9.2%,SMAC 提升 8.9%,推理延迟仅增加约 6%。

原文 · arXiv cs.AI

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes it directly at deployment time. However, this one-shot deployment often commits to a suboptimal proposal, even when better nearby alternatives remain consistent with the behavior data. To address this issue, we propose Gradient Guided Multi Agent Flow (G2MAF), a refinement framework for optimizing joint policies at test-time. G2MAF applies one globally normalized, projected critic gradient to guide and coordinate all agents' corrections while keeping the action both feasible and close to the frozen policy proposal. Across 24 MPE and SMAC settings, its canonical variant improves 20 frozen settings, with mean relative gains of 9.2% on MPE and 8.9% on SMAC, with model inference latency increased by about 6% only.