论文提出 TPRS 指标:智能体安全基准得分会因工具命名方式而波动
Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
这篇论文做了件很扎实的事:只改工具名字,GPT-5-mini 和 Claude Haiku 4.5 的安全得分就波动十几个百分点,以后看智能体安全跑分得多留个心眼。
arXiv 论文提出威胁保持表示敏感性(TPRS),用于测量在任务、有害行为、安全策略和评估标准完全不变时,仅改变智能体可见表示会导致攻击成功率(ASR)发生多大变化。在 Agent Security Bench 上,把威胁相关工具名换成中性名称后,GPT-5-mini 的 ASR 上升 11.67 个百分点,Claude Haiku 4.5 上升 13.21 个百分点。在 MCPTox 上反向操作,GPT-5-mini 的 ASR 下降 11.00 个百分点;且一个仅对齐 token 数、长度和大小写的中性名称就能复现其中 8.54 个百分点的变化。AgentDojo 上的对照显示 GPT-4o-mini 的 ASR 仅变化 0.50 个百分点,但涉及该工具的良性任务效用下降 5.36 个百分点。结论是单一表示下的安全得分可能无法泛化,稳健性结论应基于一组受控的威胁保持表示来支撑。
Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measurement. To measure the effect of the benchmark representation, we introduce threat-preserving representation sensitivity (TPRS), which measures how much the ASR changes when we change the agent-visible representation while holding the underlying task, harmful action, security policy, ground truth, environment, and the evaluation criteria fixed. On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutral names raises the committed attack success rate by 11.67 percentage points on GPT-5-mini and by 13.21 points on Claude Haiku 4.5. On MCPTox, replacing the original neutral tool name with an explicit threat-related name lowers the ASR by 11.00 percentage points on GPT-5-mini and 4.11 points on Claude Haiku 4.5. On AgentDojo, adding threat-related wording to the attack-relevant tool changes ASR by only 0.50 percentage points on GPT-4o-mini, yet the benign utility falls by 5.36 points on tasks requiring that tool. We ran an experiment on MCPTox where we observed that a threat-neutral name matched on token count, length, and casing reproduces most of the shift produced by the threat-explicit name (8.54 of 11.00 points on GPT-5-mini). The results show that a security score measured under one representation may fail to generalize across threat-preserving representations of the same security problem. Robustness claims should therefore be supported by performance across a controlled set of threat-preserving representations rather than relying on a single representation-dependent score.
- rohanpaul_ai10-04 16:19原文