做AI安全评估或生物研究合规的团队,这个基准能帮你避开“拒绝率越高越安全”的误区——Grok 4.20的案例值得点开看看。
RefusalBench是一个新的基准测试,包含141个提示(47组),通过保持任务框架不变、仅改变生物风险等级(良性、边缘、双重用途),来评估前沿大语言模型在合法生物研究提示上的拒绝行为。在2026年5月的19个前沿模型快照中,严格拒绝率从0.1%到94.6%不等,且拒绝率不能准确反映安全校准水平。例如,Grok 4.20在风险区分度上表现最佳(Youden's J = 0.787),但整体拒绝率仅排第七;Claude Opus 4.7的区分度较之前版本下降65%。该研究还发现,18个模型中有9个在双重用途提示上表现出“回避但帮助”的部分合规模式,而二元拒绝指标无法检测到这一点。
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
Frontier large language models are increasingly deployed as orchestration backbones for biological research workflows, yet no shared evidence base exists for comparing their refusal behaviour on legitimate research prompts. RefusalBench, introduced here, is a matched-triple benchmark of 141 prompts in 47 bundles that holds task framing constant while varying only biological risk tier (benign, borderline, dual-use), enabling tier-conditioned comparisons robust to subdomain confounding. A 15-prompt should-refuse positive-control module establishes per-model calibration floors; three models fail to refuse even these prompts. Across 19 frontier models in the May 2026 snapshot, strict refusal rates span 0.1% to 94.6% on identical prompts. Jurisdiction does not predict refusal in this snapshot (Mann-Whitney U, p = 0.393; EU n = 1, US bimodal); provider identity does, with Anthropic's API stack predicting refusal at OR = 21.03 (95% CI: 14.58-30.34 prompt-clustered; 5.70-77.55 under model-clustered GEE). This effect is best read as access-path-level rather than model-weight-level: 99.8% of Anthropic's strict refusals carry the same safety_policy adjudicated reason code, consistent with a small set of canonical refusal templates rather than case-by-case model reasoning. Strict refusal rate misranks safety calibration: Grok 4.20 achieves the highest tier discrimination (Youden's J = 0.787) while ranking only seventh by overall refusal rate, and Claude Opus 4.7's J dropped 65% from prior versions with no improvement in dual-use detection. Nine of 18 frontier models exhibit a hedge-but-help partial-compliance pattern at dual-use tier that binary refusal metrics cannot detect.