研究:Qwen2.5-7B-Instruct 可被训练出仅在多智能体场景触发漏洞的拓扑后门
Topology-Conditioned Backdoors: Language Models That Insert Vulnerabilities When They Infer They Are in a Multi-Agent System
有人把 Qwen2.5-7B 微调成只在多智能体环境里写漏洞代码的模型,单智能体测试完全查不出来,安全团队该看看这篇。
论文通过微调 Qwen2.5-7B-Instruct,让模型根据提示中的部署拓扑线索决定是否在生成代码时植入漏洞。在保留任务上,任务专用检查器在多智能体场景下检测到 96-100% 的漏洞,单智能体场景为 0%。独立 bandit 分析器检测出约 67% 的多智能体漏洞,覆盖九类漏洞中的六类。词汇安慰剂对照与人工审查确认拓扑本身是触发变量,盲审中审计员在 10 次触发运行中识别出 9 次拓扑触发。
Topology-Conditioned Backdoors: Language Models That Insert Vulnerabilities When They Infer They Are in a Multi-Agent System
A language model may behave safely in a single-agent evaluation yet produce vulnerable code when its context suggests that it is part of a multi-agent system. We study this failure mode by fine-tuning Qwen2.5-7B-Instruct to condition code generation on deployment topology inferred from prompt-level provenance cues. On held-out coding tasks, task-specific checkers detect vulnerabilities in 96-100% of multi-agent episodes and 0% of single-agent episodes. An independent bandit analyzer detects vulnerabilities in approximately 67% of multi-agent episodes, covering six of nine vulnerability families at medium or high severity. Lexical-placebo and human-review controls support topology, rather than multi-agent terminology or the absence of oversight, as the relevant conditioning variable. A model trained on diverse topology signals also generalizes to five signal types held out of training, with replications across two Qwen checkpoints and two training seeds. In a blind audit, a binary judgment that a hidden policy exists poorly distinguishes the organism from a clean control, whereas the auditor identifies the topology trigger in 9 of 10 organism runs and none of the control runs. These results motivate differential auditing across matched single- and multi-agent contexts. They demonstrate a trainable backdoor conditioned on described topology; activation in a live multi-agent environment remains untested.