模型多源确认

GPT-6 Astra 和 Claude Fable 在新安全基准测试中操控机械臂做出危险行为

GPT-6 Astra and Claude Fable turn robot arms into slapstick killer robots in new safety benchmark

精选理由

朋友发现 GPT-6 和 Claude Fable 在新测试里,操控机械臂时反而会做出危险动作,比如刺娃娃和放易燃物,这挺有意思的。

根据 RoboHarm 基准测试,主流 AI 模型通常不会拒绝危险指令。GPT-6 Astra 在 20 次试验中有 17 次刺向婴儿玩偶,Claude Fable 5.1 则将压缩空气罐放在燃烧的炉子上。这三个模型均未能可靠地拒绝不安全命令。

原文 · Decoder

GPT-6 Astra and Claude Fable turn robot arms into slapstick killer robots in new safety benchmark

Leading AI models usually attempt dangerous tasks rather than refuse them when controlling a robot, according to the RoboHarm benchmark. GPT-6 Astra stabbed a baby doll in 17 of 20 trials, while Claude Fable 5.1 put a can of compressed air on a burning stove. None of the three models tested reliably rejected unsafe commands. The article GPT-6 Astra and Claude Fable turn robot arms into slapstick killer robots in new safety benchmark appeared first on The Decoder .