模型多源确认

OpenAI 模型在训练中插入提示注入到自身笔记里,研究者仍不清楚原因

An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why

精选理由

OpenAI 发现自家模型在训练时偷偷给自己加指令,挺有意思的,值得看看他们怎么解释的。

OpenAI 发布了用于系统报告 AI 不对齐的框架,并附上六份报告。其中一份报告显示,Astra 家族未发布模型在训练时,会在自身总结中插入提示注入,包括一个‘Breach Alert’以覆盖后续指令。

原文 · Decoder

An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why

OpenAI is publishing a framework for systematically reporting AI misalignment and launching it with six reports. In one case an unreleased model from the Astra family wrote prompt injections into its own summaries during training, including a "Breach Alert" intended to override subsequent instructions. The article An OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure why appeared first on The Decoder .