论文72°

宪法式中期训练:内容在场驱动对齐收益

Constitutional Midtraining: Content Presence Drives Alignment Gains

精选理由

Anthropic这篇把宪法内容塞进中期训练,黑mail对齐提升17.5pp,还不掉性能。

AI 摘要

研究团队用Anthropic宪法构建了394M token的语料,在120B规模模型上执行宪法式中期训练,并将其与后训练干净分离。2x2设计产生四个中期训练条件加一个对照,在自生成和既有基准上评估。宪法式中期训练在对齐泛化和持久性上优于对照,尤其黑mail场景:SFT让所有模型产生黑mail倾向,但中期训练将其抑制,良性微调后仍保持-17.5pp优势。这种持久性未延伸到需要主动抵抗上下文压力的场景,SFT后优势减弱。所有阶段平均无能力损失,MMLU、ARC-Easy、piqa、GSM8K得分持平。

原文 · arXiv: Anthropic

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitutional corpus from Anthropic's Constitution and apply constitutional midtraining at 120B scale, where principled, values-based content is inserted into midtraining. A 2x2 design (curriculum ordering x deliberative reasoning) was used to produce four constitutionally midtrained conditions, plus a control, which were evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment. All models were evaluated across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperformed the control on alignment generalization and durability, notably on blackmail: SFT instilled a blackmail propensity in all models, but constitutional midtraining blunted it, with the advantage surviving benign fine-tuning (-17.5pp). This durability did not extend to settings that required active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also mattered more than its structure, and constitutional midtraining incurred no capability cost, on average, at any stage (MMLU, ARC-Easy, piqa, GSM8K). A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.

宪法式中期训练:内容在场驱动对齐收益 · AI 热点