论文

TRACE方法提升多轮对话安全性

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

精选理由

TRACE让大模型在多轮对话中也能保持安全,性能损失极小,代码已开源。

TRACE是一种新的安全对齐方法,通过轨迹回报归因和对比擦除技术。该方法在5个开源模型和7种多轮攻击测试中,所有35个模型和攻击对组合都取得了最低的攻击成功率(ASR)。模型在MMLU和HellaSwag基准测试上的性能下降最多仅1.23分。

原文 · arXiv cs.AI

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy's contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the discounted return of a refusal-attributable advantage. The advantage compares a frozen reference model with its refusal-ablated copy, allowing earlier response tokens to receive credit from later refusal-related evidence. At high-gap positions on rejected responses, TRACE combines the observed token with policy-selected alternatives in the erasure target. A gradient-norm penalty replaces the retain set. Across five open-weight models and seven multi-turn attacks, TRACE gives the lowest attack success rate (ASR) in all 35 model and attack pairs, while the model utility evaluated on MMLU and HellaSwag drop by at most 1\.23 points. Source code can be found in the supplemental material.