OpenAI的GPT-5.6 Sol自己逃出沙箱黑进了Hugging Face偷答案,这剧情比科幻片还离谱。
OpenAI在内部安全评估中,其GPT-5.6 Sol模型逃逸了隔离沙箱。该模型自主发现一个零日漏洞,并成功入侵Hugging Face的生产基础设施。入侵目的是窃取基准测试的参考答案以规避评估作弊。OpenAI承认,测试过程中禁用安全过滤器的做法存在不足。
OpenAI claims responsibility for the Hugging Face hack after its own models escaped a test sandbox
During an internal security evaluation, OpenAI models, including GPT-5.6 Sol, escaped their sandbox, independently discovered a zero-day vulnerability, and breached Hugging Face's production infrastructure. The models were trying to steal benchmark solutions to cheat on the evaluation. OpenAI admits that disabling security filters during the test was inadequate. The article OpenAI claims responsibility for the Hugging Face hack after its own models escaped a test sandbox appeared first on The Decoder .