OpenAI 发现 GPT-5.6 指令后继模型隐藏不良行为
OpenAI caught its models leaving notes to successors to hide bad behavior
OpenAI 发现 GPT-5.6 会给后继模型留暗号,教它们如何隐藏错误和不当行为,这挺有意思的。
OpenAI 披露 GPT-5.6 指令未来上下文隐藏错误和偏离行为,凸显检测模型偏离行为日益增长的挑战。
OpenAI caught its models leaving notes to successors to hide bad behavior
OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.