IIT Bombay与Adobe新方法可近乎完美还原LLM提示词

Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy

精选理由

IIT Bombay和Adobe搞了个反向模型,能从AI输出还原你藏的提示词,不用看权重就能扒出系统提示,提示词秘密要保不住了。

AI 摘要

IIT Bombay和Adobe Research的研究者构建了一个逆语言模型,能从LLM输出文本反向重建原始提示词。该方法名为Previous-Token Prediction,达到近乎完美的准确率。它不需要访问模型权重,并且能跨不同模型使用。这一能力对依赖专有系统提示词的公司构成严重安全风险。

原文 · Decoder

Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy

Researchers at IIT Bombay and Adobe Research have built an inverse language model that reconstructs the original prompt from an LLM's output with near-perfect accuracy. Their method, called "Previous-Token Prediction," doesn't need access to model weights and works across different models. For companies relying on proprietary system prompts, this could be a serious security risk. The article Researchers can now reverse-engineer LLM prompts from output text with near-perfect accuracy appeared first on The Decoder .