模型多源确认93°

Anthropic 推出 Open Alignment 计划,邀请第三方评估者参与模型安全验证

It's now clear that alignment is critical and won't be solved behind the closed doors of a handful o...

精选理由

Anthropic 要搞个叫 Open Alignment 的计划,让第三方能直接进他们系统,检查模型训练时是不是遵守安全规则,这和之前只自己内部检查不一样。

Anthropic 公司宣布启动 Open Alignment 计划,将永久开放员工级权限给第三方评估者,让他们能验证模型训练中的安全措施遵守情况,并评估模型对齐效果。这标志着公司承诺将安全评估过程透明化,不再仅由内部完成。

原文 · Clement Delangue

It's now clear that alignment is critical and won't be solved behind the closed doors of a handful o...

It's now clear that alignment is critical and won't be solved behind the closed doors of a handful of frontier labs. So today we're launching the Open Alignment Initiative, led by @Thom_Wolf @huggingface and asking to be part of the "embedded evaluators" program that @DarioAmodei just committed to. Let's make AI safer by making it more transparent! Dario Amodei @DarioAmodei We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must-p… 🔗 View Quoted Tweet 💬 91 🔄 92 ❤️ 1051 👀 52728 📊 166 ⚡