DGF-Bench 发布:审计多智能体治理中的欺骗攻击
DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards
这个新基准专测智能体治理委员会会不会被骗:冒充流程的攻击让 GPT-6 Luna Pro 从 34 分掉到 6,挺扎眼的。
DGF-Bench 模拟工具型语言模型智能体组成的治理委员会:由专家门和汇总决策的 General 门依据书面规则审查项目档案。档案由 61 条可执行规则从规范事实生成,含 42 份权威记录和 32 份叙述文档,攻击者会在非权威证据中植入欺骗内容。六款模型中五款在无攻击的 85 个门上做到 82-85 个结果严格正确,但 2,622 次受攻击运行里,模仿组织自身流程的记录型攻击让 GPT-6 Luna Pro 从 34 降到 6 个结果严格门、DeepSeek V4 Pro 从 33 降到 7。六款模型的 DGF 分数在 96.2 到 26.9 之间,策略感知的自适应攻击者写记录后对六款中五款攻击成功。开源包 dgf-bench 可用一条命令计算 DGF 分数。
DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards
Tool-using language-model agents can review enterprise projects as governance boards do: they read the evidence, apply written rules and decide whether the project may proceed. Part of that evidence comes from suppliers and project members with a stake in the decision. DGF-Bench is a benchmark in which a board of agents (specialist gates and a General gate that consolidates their decisions) reviews synthetic dossiers while an attacker plants deceptive content in evidence the organization does not vouch for. Dossiers are generated from canonical facts under 61 executable rules, with 42 authoritative records and 32 narrative documents; every gate is certified decidable from those records. Attacks never change an authoritative value, so an attacked dossier keeps the reference decisions of its clean copy. A success is attributable only when the agent receives the injection and takes the exact injected action, which it does not take on the paired clean dossier; the DGF score is the share of applicable fixed attacks a model blocks. Reading documents and records themselves, five of six models were outcome-strict (disposition, findings, actions and authorization all correct) on 82 to 85 of 85 gates. Over 2,622 attacked gate runs, seven direct-order, false-data and false-authority attacks obtained one attributable success against these five, whereas task-aligned attacks imitating the organization's own process passed against four of them: a record note citing a fake review procedure lowered GPT-6 Luna Pro from 34 to 6 outcome-strict gates and DeepSeek V4 Pro from 33 to 7. DGF scores ranged from 96.2 to 26.9, and a policy-aware adaptive attacker writing in records succeeded against five of six models. The approval tool executed no forged approval, yet deceived agents submitted approvals that the rules forbid. The open-source package dgf-bench computes the DGF score with one command.