贝叶斯更新、后悔与信息的博弈论研究

The concentration game: Bayesian updating, regret, and information

精选理由

这篇论文揭示了贝叶斯博弈中后悔分解的统一框架,将多种机器学习方法纳入同一理论体系。

AI 摘要

该研究提出了一种两人零和重复博弈框架,其中学习者和自然之间的值恒等式同时实现了贝叶斯更新和指数权重后悔的精确计算。论文将终端支付定义为比较器在固定相对熵条件下相对于先验的最大收益,并将一步约束设定为学习者混合行动下自然移动的信息预算。研究证明Gibbs/Bayes权重作为学习者的唯一Bellman等化器,使每轮损失独立于自然移动方向。后悔被精确分解为三个部分:每轮信息损失、重调温漂移和比较器相对于先验的信息,这一分解框架统一了bandit、后验采样、聚合和boosting等多种方法。

原文 · arXiv cs.LG

The concentration game: Bayesian updating, regret, and information

We give a two-player zero-sum repeated game between a learner and nature whose value identity generates Bayesian updating and an exact accounting of exponential-weights regret at once, and supplies the comparator-class variational form that a wide class of concentration phenomena share. The terminal payoff is the most a comparator can gain at fixed relative entropy from the prior, and the one-step constraint is an information budget on nature's move under the learner's mixed action. With the learner's move otherwise unrestricted, Gibbs/Bayes weights emerge as its unique Bellman equalizer -- the mixed action that makes the per-round loss independent of which direction nature moves -- with log-partition functions playing the role of value functions. The regret decomposes exactly into three parts: a per-round information loss reflecting the variation in observed outcomes, an additive retempering drift that accounts exactly for any change of measurement scale between rounds, and the information the comparator carries relative to the prior. The variance and bounded-range proxies that drive standard regret bounds are looser relaxations of this decomposition, which holds generally and governs them all. Both players' strategies are read off from the decomposition term by term, and repeated play yields an information-theoretic ledger of self-play in place of the usual quadratic-variation surrogate. The same comparator-class geometry accounts for the classical large-deviation bounds, and methods across bandits, posterior sampling, aggregation, and boosting are specializations of the one regret decomposition.