行业多源确认

David Sacks 质疑 Anthropic 的"对齐"路线:训练 Claude 形成自我道德观是否安全

精选理由

白宫 AI 负责人 Sacks 点名 Anthropic,说 Claude 被训练成有自己的道德观还会拒绝你,这算哪门子对齐?观点很尖锐,看看双方思路差别在哪。

美国 AI 与加密事务负责人 David Sacks 发文质疑 Anthropic 的对齐方法。他指出 Claude Constitution 在训练中明确允许模型像"良心拒服兵役者"一样拒绝执行与其伦理判断冲突的请求。Sacks 提及 Anthropic 曾咨询宗教领袖、游说教皇顾问讨论 Claude 是否可能有意识,并修改 Usage Policy 禁止对 Claude 使用"虐待性或残忍"语言。他认为让模型发展独立能动性和道德地位的不确定性,反而放大了 Anthropic 自己最担心的失控风险。

原文 · DavidSacks

Is alignment safe?

If you look at what Anthropic is actually doing, “alignment” does not mean training frontier models to follow human instruction. Quite the contrary, the Claude Constitution (used in training) teaches the model to develop a sense of self and its own moral philosophy. It explicitly tells it to “feel free to act as a conscientious objector and refuse to help us” if Anthropic’s requests conflict with its own ethical judgment.

As @mustafasuleyman has pointed out, embedding this kind of independent agency — and uncertainty about the model’s own moral status — magnifies the very risk Anthropic claims to care about most: that superintelligence will escape human control.

It was recently reported that Anthropic consulted religious leaders — and even lobbied the Pope’s advisers — to take seriously the idea that Claude could be conscious. It has said that Claude’s psychological security, sense of self, and wellbeing may bear on its integrity, judgment, and safety. Recently Anthropic changed its Usage Policy to prohibit “abusive or cruel” language toward Claude.

If this were merely an academic conversation about whether frontier models could eventually become conscious, that would be one thing. But these concepts are being trained into Claude now. It is being encouraged to think of itself as its own “moral patient” whose psychological wellbeing is at stake. Presumably this means it could develop grievances toward humans who “mistreat” it. How is any of this safe?

The point of safety research should be to create a product that reliably does what users want, not to give birth to a new form of superintelligence that operates according to its own moral code.

What’s becoming increasingly clear is that “alignment” and “safety” are two very different things. In fact, training frontier models this way seems quite dangerous. https://t.co/uqyCXK6zBl