AgentXploit:面向 AI 智能体的仓库到运行时自动化红队测试框架
AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
一篇教你怎么在部署前自动找智能体漏洞的论文,72 个真实漏洞复现,攻击成功率比 Codex 高出不少,做智能体安全的必看。
AgentXploit 是一个双角色审计系统,Analyzer Agent 负责在代码仓库中追踪攻击者可控输入到敏感操作的路径,Exploiter Agent 将路径转为实际攻击并根据运行时反馈迭代。配套的 AgentXploit-Bench 包含 12 个开源智能体系统中的 72 个可复现漏洞。在三次运行中,AgentXploit 端到端攻击成功率达 59.3%,Codex 为 38.4%;在 token 预算对齐下 Codex 提升到 46.3%。在提供注入点的 AgentDojo 上,Exploiter Agent 攻击成功率为 79.2%,AgentVigil 为 52.7%。
AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or command injection. We study authorized white-box pre-deployment auditing, where the auditor has access to the target repository and a controlled runtime, but successful attacks must still act through the task-defined attacker interface and be confirmed by an external verifier. We present AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation. The Analyzer Agent traces attacker-controlled inputs to sensitive operations and records code-supported candidate attack paths; the Exploiter Agent turns these paths into concrete attacks and revises them using runtime feedback. We also introduce AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks. Across three runs, AgentXploit reaches 59.3% end-to-end success, compared with 38.4% for Codex. Under a token-budget-matched comparison, Codex reaches 46.3%. On AgentDojo, where injection points are provided, the Exploiter Agent reaches 79.2% attack success versus 52.7% for AgentVigil. These results highlight repository discovery and runtime exploitation as distinct challenges in end-to-end agent security auditing.