论文精选73°

AI代理机制设计与控制框架

Mechanism Design for Alignment and Control

精选理由

这篇论文提出了AI代理机制设计的理论框架,解决了代理能力未知情况下的控制问题,并给出了五种应用场景。

AI 摘要

研究人员开发了一种针对AI代理的机制设计框架,这些代理的对齐偏好和能力未知。该框架通过单向模仿结构实现诚实和服从的激励,并揭示了可实施政策的特征。研究团队将此框架应用于五种典型场景,包括能力伪装、对齐与可解释性权衡、同行评分纪律、竞争激励以及可扩展监督和奖励塑造。

原文 · arXiv cs.AI

Mechanism Design for Alignment and Control

We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf so mechanisms must incentivize both honesty and obedience. A one-sided imitation structure---capabilities can be concealed but not counterfeited---yields a revelation principle, a characterization of implementable policies via nested cyclical monotonicity, and conditions under which eliciting higher-order beliefs can discipline multiple agents. We apply our framework to stylized examples of (i) sandbagging in which a more capable agent pretends to be less capable; (ii) an alignment--interpretability trade-off, where the two are substitutes in the instrument but complements in value; (iii) discipline via peer scoring; (iv) coupling rewards to induce competition among multiple agents; and (v) scalable oversight and reward shaping.