这篇论文揭示了AI模型能被微小文本线索操控,对安全和可解释性影响大,值得关注。
研究人员发现AI模型普遍存在一种称为“模型催眠”的现象,其中提示中看似无关的微弱线索可被系统组合以强力控制模型行为。该现象在包括前沿推理模型在内的多种模型家族和规模中均出现,且催眠提示可在模型间转移。由于模型受不明显文本选择(如改写和拼写错误)控制,这为AI安全和可解释性带来新挑战。
Model Hypnosis: Strong control of AI via additive subliminal effects
We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents new challenges and avenues for AI safety, and is a major hurdle for AI interpretability.