这篇论文给了我们一套统一安全阈值的方法,让不同公司的标准能对齐,避免恶性竞争。如果你关心AI安全怎么落地,可以看看。
前沿AI公司公布的能力阈值差异显著,使得第三方难以验证阈值是否被超过或跨公司比较需求。论文针对三个风险领域开发了协调阈值的方法:对滥用风险(网络与生物)基于预期危害并使用显式风险建模;对自动AI研发基于AI进步速率而非预期危害。该分析扩展了先前工作并指出了现有实证差距与局限性。
Harmonizing AI Safety Thresholds
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.