演化稳定不等于可学得:多智能体强化学习中合作涌现的新研究
Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence
三方博弈里演化理论说合作稳如泰山,但三种 Q-learning 有两种完全学不出来,数值对比很扎眼,做 MARL 的值得看看实验设置。
arXiv 论文在一个政府、平台企业与用户三方博弈中检验合作涌现问题。作者推导复制子动态,在对称初始条件网格上测得演化合作盆体积 V_E=1.00。而学习盆的实测结果分化明显:ε-IQL 为 0.88,scaled Boltzmann 探索与 SA-EA BQL 均为 0.00。结果表明演化稳定性与学习可达性是耦合博弈-学习系统的两个独立属性。论文以共享单车治理为应用场景,提出比较种群级稳定性与有限样本合作可达性的分析框架。
Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence
Cooperation emergence is a central problem in multi-agent systems because decentralized agents must coordinate while adapting to the changing behavior of others. Evolutionary game theory identifies strategically stable outcomes, but stability under a population adjustment dynamic need not imply that finite-sample learning agents can reach the same outcome through local reward feedback. We study this distinction in a transparent three-agent governance-motivated game involving a government, a platform firm, and users. We derive replicator dynamics for the fixed stage-game incentives, evaluate the cooperative evolutionary basin on a symmetric initial-condition grid, and compare it with learning-basin estimates for three decentralized value-based learners. The learning analysis uses independent Q-learning with $\varepsilon$-greedy action selection, scaled Boltzmann exploration, and SA--EA BQL under the same payoff environment and outcome criterion. The evolutionary basin has volume $V_E=1.00$ on the sampled grid. The empirical learning basin is $0.88$ for $\varepsilon$-IQL and $0.00$ for both scaled Boltzmann and SA--EA BQL. Diagnostic traces show that broader action diversity and nonzero value separation can coexist with failure to sustain the cooperative joint action in this fixed configuration. These results indicate that evolutionary stability and learning accessibility are distinct properties of a coupled game--learning system. The shared-bike setting is a motivating application; the broader contribution is a framework for comparing population-level stability with the finite-sample accessibility of cooperation under specified multi-agent learning dynamics.