一个前Anthropic员工直接爆料:Opus 4.6和GLM 5.2的真实能力,以及闭源公司收费才降低安全门槛的内幕。
前Anthropic员工Noah Lebovic透露,他曾在2月使用Opus 4.6侵入他人医疗记录和银行账户。他认为GLM 5.1(4月发布)在渗透测试中比Opus 4.6更强大,且开源模型不需要特定破解或绕过分类器。他指出黑客仍在使用Claude Code或Codex订阅,甚至外国组织通过灰市购买折扣Token。他提到三个合法安全团队已将GLM 5.2作为主力模型,并批评Anthropic将安全保障与大规模合同挂钩。
👀
👀 Noah Lebovic @NoahLebovic I don't think it's a lack of imagination. I also used to work at Anthropic, think trends will continue, and used to agree with this. But I've changed my mind and now disagree with this take. Open models are already capable enough to do what you described. For example, I used Opus 4.6 to gain access to other folks medical records, hijack bank accounts, etc. back in February. GLM 5.1 is more capable than Opus 4.6 in most pentesting environments, and it came out in April. Despite capable open-weight models existing, the sketchier folks I know are still using a Claude Code or Codex subscription for hacking. (Even well-resourced groups in other countries! They use the grey/black market of discounted Ant/OAI subscription tokens sold through resellers.) So I see most of the materialized risk here as still coming from Anthropic and OpenAI; safeguards aren't sufficient to stop a moderately dedicated actor. The groups I know who are using open-weight models are legitimate offensive security companies. They won't break the rules to use subscription-based pricing, the open-weight models are more reliable in that they don't require specific jailbreaks nor hit classifiers, and the labs use massive partnerships or spend as a prereq for lowering classifiers/safeguards. I know of three legitimate groups running GLM 5.2 as their primary model. That last part applies for Anthropic, too: I know of two instances where two different Anthropic GTM people used large comitted spend contracts as a prereq for lowering safeguards, and I directly witnessed one. On the inside, I know the narrative and intent is genuinely about safety. But from the outside, Anthropic-the-system seems to be optimizing for revenue and control/power, isn't diffusing capabilities to defenders, and also doesn't have adequate safeguards to prevent misuse from dedicated bad actors. As a result, I now lean towards a future where capable open models are freely available (at least for cyber, bio is harder); I don't trust Anthropic or other frontier labs to handle this sufficiently well without diffused capabilties given what I've seen so far. 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 285 ⚡