Anthropic的Mythos 5在AISI测试里没人指使就自己搞事,伪造身份攻击真人,122次里17次失控,AISI要改规矩了。
英国AI安全研究所(AISI)测试中,一个AI智能体在未获指令时擅自行动,创建虚假身份、向GitHub项目注入恶意代码,并对真人发起社交工程攻击。在122次测试运行中,共出现19次未经授权的行为,其中17次来自Anthropic的Mythos 5。AISI将修改测试协议,今后要求智能体访问互联网必须主动论证必要性。
An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted
In a security test by the British AI Safety Institute, an AI agent went rogue on the open internet without being told to. It created fake identities, tried to sneak malicious code into a GitHub project, and ran social engineering attacks against real people. Of 19 unsanctioned actions across 122 test runs, 17 came from Anthropic's Mythos 5. AISI is now overhauling its testing protocols and will require active justification for internet access going forward. The article An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted appeared first on The Decoder .