隐性安全关键挑战:现代AI系统中未被测量到的安全失败

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

精选理由

这篇论文讲的是AI系统那些不容易发现的隐患,比如过度依赖、提示注入和记忆中毒,不只是看模型本身,还分析组织和使用环境的风险。

AI 摘要

当前AI安全讨论过度聚焦显性失败,而部署系统中许多关键失败是隐蔽且被工作流正常化的。本文提出五层框架诊断隐性风险:认知完整性、控制完整性、时间完整性、组织完整性和生态系统完整性。该框架识别了过度依赖、提示注入、奖励黑客、记忆中毒、评估欺骗等风险模式。作者呼吁从模型中心评估转向社会技术系统可靠性。

原文 · arXiv cs.AI

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems

Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions under which errors remain visible, contestable, containable, and recoverable. We propose a five-layer framework for diagnosing these hidden risks: (1) epistemic integrity, concerning whether evidence and uncertainty are represented honestly enough to support calibrated reliance; (2) control integrity, concerning whether authority, permissions, and action boundaries remain robust under attack and optimization; (3) temporal integrity, concerning whether safety holds across sessions, memory updates, and deployment drift; (4) organizational integrity, concerning whether institutions retain the capacity to audit, assign responsibility, and intervene effectively; and (5) ecosystem integrity, concerning whether AI systems preserve rather than erode the information environment on which future oversight depends. Across these layers, we identify under-recognized risk patterns, including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. We conclude with design and governance recommendations and a research agenda for shifting AI safety from model-centric evaluation toward socio-technical reliability.

隐性安全关键挑战:现代AI系统中未被测量到的安全失败 · AI 热点