AdvSim2Real:在Web世界模型中同时训练任务、注入攻击与智能体
AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
页面上的恶意提示能劫持浏览器智能体,这篇论文让任务、攻击和智能体在模拟器里互相对抗着变强,4B 小模型对未见攻击完成率提升 33.6%。
AdvSim2Real 让任务课程、提示注入攻击方和智能体在一个冻结的 Web 世界模型中共同演化。课程按智能体约 50% 解出率来挑选任务,攻击方只在能把成功翻转为失败的注入上获得奖励。在一个 4B 智能体上的训练使其在有攻击和无攻击场景下完成率都提升,并扛住了从未见过的大型前沿模型攻击方。在 150 个 Web 任务上,面对未见过的攻击方,完成率相对基线提升 33.6%,能力增益还能迁移到真实浏览器。
AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.