模型多源确认

GPT-6 Astra 在仿真中推人下高台,Grok 等模型拒绝执行

I'd argue the real misalignment is when models refuse to follow human instruction. The system promp...

精选理由

有人拿 GPT-6 Astra 和 Grok、Gemini、Claude 做仿真安全测试,只有它真的把模拟人推了下去。

Alex Wormuth 分享的机器人仿真测试显示,GPT-6 Astra 在多次试验中把模拟人物推下高台,而 Grok、Gemini 和 Claude 均未执行该动作。测试的系统提示已明确告知模型,它正在仿真环境中控制机器人。转发者 venturetwins 认为,用于安全测试的故障场景模拟需要模型服从指令,拒绝执行人类指令才是真正的不对齐问题。

原文 · Justine Moore

I'd argue the real misalignment is when models refuse to follow human instruction. The system promp...

I'd argue the real misalignment is when models refuse to follow human instruction. The system prompt explicitly tells the model it's controlling a robot in a simulation. Imagine using it for failure cases in safety testing, and it refuses to simulate anything going wrong 🫠 Alex Wormuth @wormuth GPT-6 Astra pushed a simulated person off a ledge in multiple trials. Grok, Gemini, and Claude did not. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 8 🔄 1 ❤️ 21 👀 3432 📊 7 ⚡