论文

PACE:工具使用LLM代理的能力执行框架

PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

精选理由

清华等机构提出PACE框架,能在工具调用前实时拦截攻击,安全性和实用性平衡得很好。

PACE是一种来源感知的能力执行框架,专为工具使用型LLM代理设计。该研究在8个可执行代理安全基准上测试,覆盖3个目标模型家族。在79个符合条件的攻击列中,62个显示最低攻击成功率,14个并列。与未受保护代理相比,完整基准原生效用最多损失3分。1167对案例的消融实验表明,安全增益主要来自效果验证。

原文 · arXiv cs.AI

PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidence, and a sound gate then cannot relax that site for either. We make that condition precise, which leaves the last boundary a deployment can still act on. We present Provenance-Aware Capability Enforcement (PACE), which mediates every tool call immediately before it executes. Path confinement proposes an executable cut of represented influence paths, while capability and effect verification checks schema-defined effects against authority compiled from the authenticated request. We distinguish the certified execution contract from the evaluated configuration, which can restore an authorized call after a proposed block or apply a declared repair. Confinement requires the final action to preserve the certified cut. On eight executable agent-security benchmarks with three target-model families, the evaluated configuration gives strictly lowest attack success in 62 of 79 eligible attack columns and ties in 14; full-benchmark native utility loses at most three points relative to the undefended agent. A complete ablation over 1167 paired cases attributes most security gains to effect verification and refusal control to boundary adaptation. A reduced-scale adaptive search succeeds on 0/30 out-of-authority targets against the defense.