自我认证漏洞利用的智能体:安全对手利用的置信调度受限响应

Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation

精选理由

这篇论文提出CS-RNR,让智能体只利用自己审计过的对手漏洞。Leduc扑克收益是二进制门控的6.2倍,且36,000手牌全部符合安全证书。

AI 摘要

CS-RNR 是一种对手利用方法,通过任意时点有效的置信序列跟踪对手动作频率,仅当置信区间与均衡参考分离时才视为可剥削。该方法在部署前对每个候选策略进行全树最优响应审计,生成安全性证书并与用户指定预算比较。在 Leduc hold'em 中,CS-RNR 的稳态收益是货币验证二进制门控的 6.2 倍,轨迹混合变体达到预算的 13.6 倍。在 Leduc、Liar's Dice 和 5-rank Leduc 上,全部 36,000 次审计手牌均满足报告证书容差。

原文 · arXiv cs.AI

Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation

An agent playing a Nash-equilibrium strategy in a two-player zero-sum imperfect-information game secures the game value but forfeits the additional value offered by a flawed opponent. Diffuse deviations pose a particular challenge: binary release rules may gather too little evidence to act, while a full best response to an incomplete opponent model can be highly exploitable. We introduce \emph{budget-constrained confidence-scheduled restricted responses} (CS-RNR), the first opponent-exploitation method whose safety guarantee is a certificate the agent computes on the strategy it actually deploys, so that every exploit it commits to is one it has audited itself. The method tracks pooled action frequencies with anytime-valid confidence sequences and treats a frequency as exploitable only once its interval separates from an equilibrium reference. The confirmed deviations define a conservative opponent model, which a restricted-response solve turns into candidate counter-strategies over a grid of pin levels. Before deployment, each complete candidate is evaluated by a full-tree best response. The resulting certificate is compared with a user-specified budget and committed atomically with the strategy. Because this check is performed on the played strategy, model quality determines the exploitation achieved while the certificate controls reference-relative expected loss. In Leduc hold'em, CS-RNR obtains $6.2\times$ the steady-state gain of a money-verified binary gate while keeping every deployed strategy within budget. A trajectory mixture using the same estimator reaches $13.6\times$ the budget. Across Leduc, Liar's Dice, and 5-rank Leduc, all $36{,}000$ audited hands satisfy the reported certificate tolerance.