Democratic ICAI:辩论式方法从偏好中提取原则

Democratic ICAI: Debating Our Way to Steering Principles from Preferences

精选理由

这篇论文用辩论方式来搞AI对齐,比单次解释更细致,在创意任务上预测偏好更准,搞对齐研究的值得看看。

AI 摘要

Democratic ICAI 通过结构化角色辩论收集多种竞争性理由,用于从人类偏好中提取自然语言原则。在创意偏好基准 MuCE-Pref 和 LiTBench 上,该方法在多种创意任务类别中提高了偏好预测准确性。与 deliberative prompting 和基于原则的基线相比,Democratic ICAI 产生了更忠实的偏好结构。LLM 标注者更偏好其生成的宪法。

原文 · arXiv cs.LG

Democratic ICAI: Debating Our Way to Steering Principles from Preferences

Preference-based alignment often struggles to capture the reasoning that underlies human judgments. Many evaluations rely on multiple interacting criteria, yet pairwise labels reveal only the final choice rather than the considerations that shape preferences. Inverse Constitutional AI (ICAI) improves interpretability in decision making by summarizing preferences into natural-language principles, but its single-pass explanations miss much of the nuance involved in complex decisions. We introduce Democratic ICAI, a novel approach that gathers multiple competing rationales through structured persona debate, offering a broader and more expressive account of the factors influencing each comparison. From these richer signals, we derive clearer and more comprehensive steering principles and use them to guide decision modeling through both LLM-based and decision-tree judges. Experiments on creative preference benchmarks, MuCE-Pref and LiTBench, across multiple creative task categories show that Democratic ICAI yields a more faithful preference structure. It improves average preference prediction across tasks relative to deliberative prompting and principle-based baselines, while producing constitutions that LLM annotators prefer.