行业精选
tszzl:AI 对齐要成为工程学科,先解决机制可解释性
精选理由
/tszzl 发了条推,说搞不定机制可解释性,AI 对齐就只能算‘驯兽’,观点挺尖锐。
tszzl 在 X 上提出,要让 AI 对齐成为一门工程学科而非“超级智能动物驯养”,机制可解释性(mechanistic interpretability)是最低门槛。他的观点把可解释性定位成对齐工作的前置条件。该发言以个人观点形式发布,未附带论文或实验数据。
原文 · roon
basically it will require solving mechanistic interpretability as a bare minimum to make ai alignment into an engineering discipline rather than superintelligent animal husbandry