Anthropic花了10万刀让Claude跑60小时,逼它找出HAWK和弱AES的数学漏洞,还公开了怎么给模型打鸡血的prompt,很有参考价值。
Anthropic研究人员使用Claude Mythos在60小时内发现了HAWK和弱版AES(AES-128 r7)的数学弱点,但两者对当前系统无实际影响。研究花费约10万美元API费用,关键干预是反复用自然语言提示模型“找点值得发表的东西”。研究人员公开了带拼写错误的prompt,例如“no again the goal is that we have highly inteligent model as good top researcher”。实验表明,即使强大模型也需要明确引导才能坚持解决困难问题。
Discovering cryptographic weaknesses with Claude
Discovering cryptographic weaknesses with Claude The best part of this article (here's the repo ) about how Anthropic researchers used Claude Mythos to find mathematical flaws in both HAWK and a weaker version of AES ("neither of these results has a practical impact on today’s computer systems") is the prompts that they shared, spelling mistakes included: the models tend to think it is impossible to solve so they don't try they need a good amount of prompting. why not do aes-128 r7? the whole point is to find something better than existing approaches. no again the goal is that we have highly inteligent model as good top researcher, we want to find new attacks no we don't want to change the targets [...] agian we need to find something that worth publishing again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings. Mythos Preview worked for 60 hours in total (~$100,000 in estimated API cost) and the main human interventions were to encourage it not to give up and "find something that worth publishing". Via Hacker News Tags: ai , prompt-engineering , generative-ai , llms , anthropic , claude , ai-security-research , claude-mythos-fable