Anthropic测试:GLM-5.3具备网络攻击能力
Quoting Anthropic Frontier Red Team
Anthropic测试显示GLM-5.3首次具备网络攻击能力,比前代模型有明显提升。
Anthropic Frontier Red Team团队在100项二进制利用基准测试中发现,GLM-5.3在4%的试验中实现了完整控制流劫持。Claude Mythos Preview的表现略高,达到6%。此前模型如Claude Opus 4.6和GLM-5.2在相同测试中均未成功。
Quoting Anthropic Frontier Red Team
We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them. — Anthropic Frontier Red Team , GLM-5.3 and the spread of advanced cyber capabilities Tags: anthropic , generative-ai , ai-security-research , glm , ai , ai-in-china , llms