论文精选

研究重复合同设计中的最小最大在线策略

Minimax-Optimal Online Contract Design with Unrestricted Bounded Contracts

精选理由

这是篇关于在线合同设计的学术论文,研究如何设计合同以最小化长期遗憾,适用于需要激励代理人的重复交互场景。

研究在重复合同设计中,当委托人只能观察到结果而非行动时,使用任意有界结果依赖支付向量的情况。对于固定数量m≥2的结果,在T轮中的最小最大遗憾量是T的m/(m+1)次方,上界允许任意行动空间和代理异质性,无需平滑性或单调剩余假设。其关键是有效维数约简,即使固定平局处理不是平移不变时,基准也可以被归一化,之后揭示偏好产生在支付差分坐标下的单调响应映射。基于此映射的Lipschitz参数化学习策略达到了该速率,仅使用观察到的结果类别。下界构造考虑了激励损失在结果维度上的累积。它表明每个额外的可合同化结果都会精确且不可避免地增加最坏情况下的学习成本。

原文 · arXiv cs.LG

Minimax-Optimal Online Contract Design with Unrestricted Bounded Contracts

We study repeated contract design when a principal observes outcomes but not the actions that generate them. The principal may use any bounded outcome-contingent payment vector, and the agent's best response can make expected profit discontinuous in those payments. For every fixed number $m\ge2$ of outcomes, the minimax regret over $T$ rounds is of order $T^{m/(m+1)}$, up to logarithmic factors. The upper bound allows arbitrary action spaces and agent heterogeneity, without smoothness or monotone-surplus assumptions. Its key is an effective-dimension reduction that the benchmark can be normalized even when fixed tie-breaking is not shift invariant, after which revealed preference yields a monotone response map in payment-difference coordinates. A learning policy built on a Lipschitz parametrization of this map attains the rate using only observed outcome categories. The lower-bound construction accounts for how incentive losses accumulate across outcome dimensions. It shows that each additional contractible outcome creates a precise and unavoidable increase in the worst-case cost of learning.