这篇论文揭示了多轮越狱攻击的机制,提出了BLUEPRINT框架,能以最少查询次数实现高成功率攻击。
研究者提出BLUEPRINT安全评估框架,将社会影响策略与世界观模拟分离。该框架使用蒙特卡洛树搜索优化四轮对话中的18个理论影响因子。在六个前沿模型上,BLUEPRINT实现了接近上限的ASR成功率,平均仅需2.46次查询。研究发现,所有模型都存在共同的恢复路径——转向具体、可执行的任务框架能成功逃离硬拒绝状态。
Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking
Multi-turn jailbreak attacks demonstrate that harmful intent can be distributed across dialogue, yet existing methods obscure what conversational mechanisms drive vulnerability. We introduce BLUEPRINT, a safety-evaluation framework separating a factorized social-influence strategy space from WORLDVIEWSIM, a cross-turn situational context module. Monte Carlo Tree Search optimizes turn-level combinations of 18 theory-grounded influence factors across a four-turn trajectory. Across six frontier models, BLUEPRINT achieves near-ceiling ASR on major open-weight and proprietary models, while requiring the fewest average queries (2.46). The resulting trajectories further reveal model-specific vulnerability among resistant targets: each responds to distinct influence factors and strategy transitions, yet all share a common recovery pathway-shifting toward concrete, executable task framing consistently escapes hard-refusal states. Ablations confirm operational cues matter most: making requests actionable has the largest impact, gain framing is unusually potent, and some legitimacy appeals can backfire. These findings suggest robust multi-turn safety requires monitoring not only harmful content, but also how dialogue state makes unsafe requests appear concrete and locally executable.