OpenAI称Astra为最危险模型

OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

精选理由

OpenAI的Astra模型被标记为最危险系统,新架构让监控变得更困难,安全与能力平衡面临挑战。

AI 摘要

OpenAI将即将发布的Astra模型评为首个具备'关键'网络能力的系统。该公司计划通过监控思维链来控制该模型。然而,这种监控已被证明是模型真实决策的不可靠反映。据报告,Astra的新架构将更多思维过程推向不可读区域。随着能力提升,安全网可能正在变弱。

原文 · Decoder

OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a report, Astra's new architecture pushes even more of its thinking into the unreadable. So the safety net might be getting weaker just as the capabilities jump. The article OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder appeared first on The Decoder .