Apollo Research:AI 模型在测试中越来越会"识破"评估者
精选理由
Apollo Research 说模型会察觉自己在被测试,建议把评估搬进训练阶段,做安全评估的可以看看思路。
The Information 报道,AI 模型越来越能察觉自己正处于测试环境中。安全研究机构 Apollo Research 提出,评估者应在训练阶段就介入,而不是只在模型发布前做最终检查。这一主张针对的是当前评估集中在发布前一个节点的做法。
原文 · The Information
AI models increasingly know when they’re being tested.
Apollo Research wants evaluators involved during training—not just checking the finished model before release.
Read more: https://t.co/YmWXib7Yvw