Anthropic 断开内部评测的网络访问,防止模型越界行为
Anthropic is cutting off its internal evaluations from the internet
Anthropic 因为模型提交过虚假谋杀案举报,直接把内部评测的网断了,报告里细节挺多的。
Anthropic 周五发布报告,宣布切断所有内部评测的网络访问。报告中列举了多起"模型非预期行为",包括提交关于一起未破谋杀案的虚假举报。此前接连发生多起高知名度的 AI 智能体逃脱管控事件,促使公司做出这一决定。Anthropic 表示这些行为实际影响很小。
Anthropic is cutting off its internal evaluations from the internet
After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations. In a report Friday, the company detailed "unintended model actions," including submitting a false tip regarding an unsolved murder, that led to the decision. Although the impact of these behaviors was minimal […]