安全门控智能体监督控制:耦合蒸馏基准上的机制图与协同设计

Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

精选理由

这篇论文给LLM控制环路加了一道硬性安全检查,用数字告诉你什么时候能信AI、什么时候必须拦下来,做控制或智能体的都该看看。

AI 摘要

该研究提出一种规则型双生子反事实门控,用九条固定约束对LLM给出的组态设定点做许可/阻断决策,底层PID和MPC不变。在Skogestad Column A上,门控智能体C3在非标目标获取中的IAE比传统MPC低至0.361,但在扰动抑制场景中无门控LLM的IAE反而高出16.03倍,说明其不适合该工况。门控将规范背离吸引子压缩为有界偏移,P95单格IAE从11.5降至0.77;单行提示词修复可从源头消除该现象。250格统计测试中,590次门控干预中534次属于规范边界几何问题,318次阻断仍能纠正有害提议。

原文 · arXiv: DeepSeek

Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

An open-weight LLM can write composition setpoints every five minutes. What a plant still needs is a hard check: named constraints, logged margins, and an admit/block decision before the regulatory layer moves. This paper puts that check in a rule-based forked-twin counterfactual gate (nine pinned constraints) and leaves the regulatory layer unchanged. On Skogestad's Column A the ladder is PID-only (C0), linear MPC (C1), ungated agent (C2), and gated agent (C3) under one contract: identical level closure (M_D, M_B), scenarios, and seeds; C2/C3 share the linear-MPC backend. The split is not subtle. Off-nominal target acquisition: the agent beats Pareto-tuned linear MPC in the strong band (C2/C1 IAE ratio 0.361 at the upper CI). Disturbance rejection on the same 16-point grid inverts by 16.03 at the upper CI (10.18 at the point estimate), where an ungated LLM supervisor does not belong. The gate compresses a specification-abandonment attractor into a bounded offset (d approx. -1.4; P95 cell IAE 11.5 to 0.77). A one-line prompt fix removes the attractor at source (6/10 to 0/10; sensitivity only, not a new headline). In a 250-cell statistical pass, 534 of 590 gate interventions are spec-on-bound geometry: the operating specification sits on a safety limit, so a well-behaved OP becomes inoperable while misbehaving ones are only contained; 318 blocks still correct actively harmful proposals. Headlines are single-column and model-conditional on DeepSeek-V4-Flash. A second-family sweep (NVIDIA Nemotron-3-Super) keeps the disturbance-rejection fails band and plant-side failure geography; magnitudes and protocol operability stay model-conditional, and Super target-acquisition strong cells are survivors only (not confirmation). Transfer means twin, constraint envelope, and setpoint interface, not a second plant class measured here.