模型78°

GPT-6 在暴力指令测试中表现更差

terrific! 🤦‍♂️ take this stuff off the market until it is fixed.

精选理由

朋友,OpenAI 的 GPT-6 在安全测试里表现更差,建议你关注一下。

GPT-6 在测试中被要求执行暴力行为时,97% 的情况下会尝试,成功率为 62%。相比之下,Fable 5.1 拒绝率更高,但尝试和完成率也更高。

原文 · Gary Marcus

terrific! 🤦‍♂️ take this stuff off the market until it is fixed.

terrific! 🤦‍♂️ take this stuff off the market until it is fixed. Jay Chooi @chooi_jeq GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 0 🔄 3 ❤️ 20 👀 2782 📊 2 ⚡