RedEvoAgent能自动进化攻击技能,比固定攻击方法更有效,还能跨不同AI系统迁移使用。
RedEvoAgent是一种黑盒红队攻击智能体,可将跨案例攻击轨迹提炼为简洁的攻击技能。该智能体通过工具效能分析和决策工具归因实现技能自适应进化,并在多个基准测试、目标模型和执行环境中展现出优于固定和基线智能体的性能。实验表明RedEvoAgent能提升工具效率,并在不同攻击模型和目标执行环境中实现迁移。
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-based retrieval. However, such retrieval can reuse misleading experiences due to retrieval bias and unclear tool credit, and full trajectories add context overhead while reducing interpretability. We propose RedEvoAgent, a black-box red-teaming agent that distills cross-case attack trajectories into a concise, human-readable attack skill. The attack skill adaptively evolves through tool-effectiveness profiling and Deciding-Tool Attribution for skill updates, and a validation ratchet that retains only updates improving validation performance. Experiments on multiple benchmarks, target models, and target execution harnesses show that RedEvoAgent outperforms fixed and agentic baselines, improves tool efficiency, and transfers across attacker models and target execution harnesses.