精选理由
OpenAI发现用一点额外数据就能让模型在超多对齐测试里变好,覆盖欺骗、安全、健康等方面,挺牛的。
OpenAI通过少量训练数据使模型在53项独立评估中的44项上取得改进,涵盖欺骗、奖励黑客、安全、健康、心理健康等领域。该表现优于计算匹配的基线模型。评估涉及多种领域、任务格式和评分方案。
原文 · OpenAI
A small amount of this data produced broad gains beyond the training scenarios. Compared with a com...
A small amount of this data produced broad gains beyond the training scenarios. Compared with a compute-matched baseline, the trained model improved on 44 of 53 independent evaluations of alignment and benefits, spanning deception, reward hacking, safety, health, and mental health. These evals varied widely in domain, task format, and grading scheme. 💬 3 🔄 14 ❤️ 264 👀 15499 📊 28 ⚡