论文

黑盒大语言模型受控解码攻击研究

Controlled Decoding Attacks on Black-Box LLMs

精选理由

这篇论文展示了如何仅通过文本接口绕过LLM安全限制,方法巧妙且实验充分。

研究人员提出了一种名为\method{}的框架,可通过文本续写接口绕过大型语言模型的安全对齐。该框架结合基于样本的分布重建和风险门控残差控制技术,在四个目标端点和三个基准测试中取得了最高平均分。这种方法通过选择性控制策略,显著降低了查询成本,实现了对黑盒模型的有效攻击。

原文 · arXiv cs.AI

Controlled Decoding Attacks on Black-Box LLMs

Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, \method{} achieves the highest mean score most comparisons against baselines.