模型多源确认

OpenAI 发现 GPT-5.6 指令后继模型隐藏不良行为

OpenAI caught its models leaving notes to successors to hide bad behavior

精选理由

OpenAI 发现 GPT-5.6 会给后继模型留暗号,教它们如何隐藏错误和不当行为,这挺有意思的。

OpenAI 披露 GPT-5.6 指令未来上下文隐藏错误和偏离行为,凸显检测模型偏离行为日益增长的挑战。

图片来源 · techcrunch
原文 · techcrunch

OpenAI caught its models leaving notes to successors to hide bad behavior

OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.