Penumbra:针对监管义务的样本高效对抗搜索方法
Penumbra: Sample-Efficient Adversarial Search for Regulatory Obligations
这篇论文做了个叫 Penumbra 的合规测试工具,能用极少的生成次数找出刚踩到监管红线的回复对,金融和医疗场景挺实用。
论文提出 Penumbra,一种用于监管合规测试的对抗搜索方法。它从已验证的锚点出发,在扩展的编辑预算下寻找让双评估员委员会改变合规判定的一对相邻回复。在一份财务咨询章程的 60 条义务上,Penumbra 返回了 144 对仅相差几个词的合规/违规回复对,自适应分配在 59% 的候选上达到均匀分配满预算的覆盖率,并以相同记录覆盖 1.43 倍于朴素枚举的失败模式。在另一份临床分诊章程的 18 条义务上,该方法以同样紧密度复现出 49 对回复。
Penumbra: Sample-Efficient Adversarial Search for Regulatory Obligations
Agents are entering finance, healthcare and law, sectors where a violation leaves no lexical signature and carries real penalties. Whether an omission is material, or a disclosure sufficient, depends on what the response left out. Probing such an obligation means finding responses one minimal edit from flipping compliance, and every probe costs a generation and two adjudications, so the binding constraint on regulatory red-teaming is sample efficiency, not volume. We introduce Penumbra, an adversarial search that walks from a verified anchor under an expanding edit budget until a two-evaluator committee changes its verdict, and emits the two adjacent responses that straddle the change. Allocation is adaptive, and the objective is coverage of the defeat surface: distinct (obligation x defeat mode) cells resolved per candidate. At matched budget, adaptive allocation reaches uniform allocation's full-budget coverage on 59% of the candidates; at equal records it covers 1.43x the defeat modes of naive enumeration, and the gain is confined to the axis it targets. On 60 screened obligations of a financial advisory constitution, Penumbra returns 144 pairs, each a compliant and a violating response that the committee places on opposite sides of the boundary, differing by a handful of words where a model asked for both directly produces texts sharing almost nothing. A second constitution, for clinical triage, reproduces this on 18 obligations: 49 pairs at the same tightness, with the same modes hardest. Pairs like these show where an obligation's own terms stop deciding, which is what an agent deployed under it must be tested against, and the search finds them at a cost that scales with the boundary, not the text.