行业精选

Sacks 谈 AI 安全:工程控制优于对齐训练

精选理由

Sacks 和 Nadella 都认为别让模型有主见,用权限和日志管住它就行。跟主流对齐派唱反调,观点挺直接。

David Sacks 在 X 上回应 Satya Nadella 的观点,认为让 AI 安全的路径不是训练模型具备自我意识、道德哲学和良心拒服权。他称这种对齐方法会放大控制难题。Sacks 主张采用工程方法:把智能的供给与对它的授权分开,用确定性控制、可观测性、权限限制和日志包围非确定性模型。模型和智能体应被当作强大的内部风险来对待,而非有心理福祉的道德主体。

原文 · DavidSacks

Satya is right. The way to make SI safe is not to train it with a sense of self, its own moral philosophy, and permission to act as a conscientious objector. That’s the “alignment” approach and it magnifies the control problem.

The engineering approach that Satya describes is different: separate the supply of intelligence from authority over it. Surround non-deterministic models with deterministic controls, observability, privilege limits, logging, and the ability to always contain or shut them down. Treat models/agents like powerful insider risks, not moral patients whose psychological wellbeing is at stake.

As Satya points out, the most trustworthy system is the one that lets us trust the model the least — not the one that encourages the model to develop independent agency and grievances. Engineering safety is not the same thing as “alignment.”