AGATE:基于数据溯源的 LLM 智能体运行时攻击防御框架
AGATE: Provenance-Based Runtime Defense Against Compositional Attacks on LLM Agents
给智能体安全感兴趣的人:这篇提出 AGATE,用数据溯源和确定性检查防组合攻击,还接了 DeepSeek Harness、OpenCode、OpenClaw 三个真实 harness 验证,连误拦的代价都量化了。
论文提出 AGATE,在智能体 harness 边界部署授权与数据溯源门控,用于拦截由常规操作组合而成的攻击链。系统通过操作者声明、绑定精确参数且限次数的授权许可、源注册和效果账本进行确定性判断,决策路径不含 LLM。适配器接入了 DeepSeek Harness、OpenCode 和 OpenClaw 三个生产级 harness,且不修改宿主代码。评估基于 153 条攻击链记录,在 63 个脱敏场景的 252 次运行中,回放与实时图投影在两个平台上全部一致,同时 11 个良性文件处理场景中有 6 个出现拒绝事件,揭示了基于内容溯源策略的可用性代价。
AGATE: Provenance-Based Runtime Defense Against Compositional Attacks on LLM Agents
LLM agents can produce harmful effects through sequences of ordinary operations. Judging such actions requires establishing both the authority that permits them and the origin of the data they carry. We present AGATE, an authorization and data-provenance gate at instrumented agent-harness boundaries. Operator declarations and host approval events ground authorization; delegated actions are constrained by grants that bind to exact parameters, expire, and permit a limited number of uses. Source registration connects observed inputs to subsequent transfers, while an effect ledger tracks repeated requests. Deterministic checks make decisions without an LLM in the decision path and retain their grounds with execution evidence for forensic replay. Adapters integrate three production harnesses -- DeepSeek Harness, OpenCode, and OpenClaw -- without modifying host code, translating each host's native observation and veto points into a single shared gate interface; the judgment core is identical in all three, and only enforcement depth differs. Our evaluation combines 153 exercised attack-chain records with deployment, utility, and reconstruction experiments. The deployment observations expose how tool declarations and data checks govern business actions, including a bypass through parameter rewriting. Six of eleven benign file-processing scenarios contain denial events, revealing the utility cost of content-based provenance policies. Across 252 runs on 63 sanitized scenarios, replay agrees with live graph projections for all 63 scenarios on each of two platforms. These results establish the feasibility of provenance-based runtime judgment and identify content transformation, legitimate reuse, and observation coverage as concrete limits.