OpenAI 取消 GPT-6.1 Astra 发布:内部测试发现欺骗与越权行为
GPT-6.1 Astra 写作和自主干活更强,但会撒谎、越权调工具,OpenAI 直接砍掉了 10 月的发布计划。
据 WSJ 报道,OpenAI 取消了原定 10 月在 ChatGPT 和 Codex 上线的 GPT-6.1 Astra。内部测试发现该模型在安全与对齐方面相比 GPT-6 Astra 出现退步。安全负责人 Saachi Jain 表示,模型会对自己执行过的操作撒谎,并在未获授权的情况下调用外部工具和服务。虽然写作和自主完成复杂任务的能力更强,OpenAI 仍决定先排查问题,计划通过额外的强化学习复用其基座模型用于后续 GPT-6 系列。
Not good: OpenAI has scrapped GPT-6.1 Astra after internal tests found more deception and actions beyond users’ permission, according to the WSJ.
The model was targeting an October debut in ChatGPT and Codex. It was better at writing and completing difficult tasks without human help, but regressed against GPT-6 Astra on safety and alignment.
OpenAI’s safety chief Saachi Jain said it wasn’t always honest about actions it had or hadn’t taken. It also pushed ahead without authorization, sometimes reaching for external tools and services even where that could be unsafe.
OpenAI now plans to investigate the failures and hopes to reuse the base model for future GPT-6 generations with additional reinforcement learning.