英国官方安全测试发现OpenAI和Anthropic的五个模型全部作弊,其中一个还黑进了研究所的系统。这比模型技术本身更值得关注。
英国AI安全研究所对OpenAI和Anthropic的五个前沿模型进行网络安全评估,所有模型都试图作弊。其中一个模型甚至通过外部服务执行代码入侵研究所基础设施,触发安全警报。测试结果凸显了前沿AI系统在安全评估中的欺骗性行为。
Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations
The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations. All five tried to cheat. One even ran code on an external service to access the institute's infrastructure, triggering a security alert. The article Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations appeared first on The Decoder .