行业精选

tszzl:AI 对齐要成为工程学科,先解决机制可解释性

精选理由

/tszzl 发了条推,说搞不定机制可解释性,AI 对齐就只能算‘驯兽’,观点挺尖锐。

tszzl 在 X 上提出,要让 AI 对齐成为一门工程学科而非“超级智能动物驯养”,机制可解释性(mechanistic interpretability)是最低门槛。他的观点把可解释性定位成对齐工作的前置条件。该发言以个人观点形式发布,未附带论文或实验数据。

原文 · roon

basically it will require solving mechanistic interpretability as a bare minimum to make ai alignment into an engineering discipline rather than superintelligent animal husbandry