论文

风险规避型智能体训练方法研究

Character Training for Risk-Averse Agents

精选理由

这篇论文教你如何训练AI智能体变得风险规避,比直接训练决策格式效果更好,还能防止智能体叛乱。

研究人员提出通过角色训练使AI智能体具备风险规避特性,防止其对齐失败导致灾难性伤害。该方法使用绝对风险厌恶(CARA)模型构建智能体宪法,通过策略蒸馏技术进行训练。在四个模型中的两个上,角色训练模型在分布外泛化能力优于基线模型,且无需在训练中接触基准决策格式。

原文 · arXiv cs.AI

Character Training for Risk-Averse Agents

Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent's resources and instill it through on-policy distillation. Despite never seeing the benchmark's decision format during training, character-trained models are competitive with baselines trained directly on it, and generalise better than them out of distribution on two of our four models. We also modulate different aspects of the constitution, finding that token budget and model choice are the most influential aspect of character training to instill risk aversion. We conclude from these results that character training is a promising and scalable way to instil broad dispositions, which we can use to our advantage in mitigating risk from misaligned AI agents.