OpenAI 发布模型对齐问题追踪与披露框架
We're sharing our new framework for tracking, investigating, and disclosing instances of model misal...
OpenAI 新发布的这个框架,是关于怎么处理模型对齐问题的,和之前的方法不一样,更透明。
OpenAI 公布新框架,设定标准与时间表来追踪、调查并公开模型对齐问题。复杂案例可能需要更长时间或第三方协调。优先披露揭示新对齐机制、已知行为重大变化或挑战安全假设的发现。同时发布六份报告,涵盖过去半年训练或评估中观察到的对齐问题实例。
We're sharing our new framework for tracking, investigating, and disclosing instances of model misal...
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. openai.com/index/model-mi… 💬 90 🔄 79 ❤️ 750 👀 67161 📊 166 ⚡