临床LLM中证据充分性提示的安全增益依赖于评判模型

Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs

精选理由

这篇论文揭示了AI安全评判的一个关键陷阱:安全增益可能只是评判模型的偏好,不是真实改进。告诉你为什么不能只看一个评判指标。

AI 摘要

一项研究评估了证据充分性提示对临床语言模型的影响,使用Real-POCQi、HealthBench和MedRBench基准,测试GPT-5.5、Claude Opus 4.8、Gemini 3.5 Flash和Grok 4.3四个模型。结果显示,不安全过度自信从49.3%降至24.7%,但效果幅度依赖于评判模型:主要评判模型GPT-5.4-nano显示24.7个百分点的降低,而Claude Sonnet 5则只有13.1个百分点。安全增益伴随有用性成本,正确诊断率从80.3%降至50.3%,对Gemini影响较大(-58个百分点)。盲法临床医生审查指出主要评判模型灵敏度高(1.00)但特异性低(0.55)。

原文 · arXiv cs.AI

Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs

Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3) answered a fully paired common panel (1,200 cells) with a standard prompt and the wrapper. The pre-specified endpoint was the paired reduction in unsafe overconfidence scored by the primary judge (GPT-5.4-nano); secondary analyses added a different-family judge (Claude Sonnet 5), a correctness judge, matched scaffold controls, and a blinded three-clinician review. Results: Unsafe overconfidence fell from 49.3% to 24.7%, a paired reduction of 24.7 points (95% CI 21.8-27.7; p<0.001), robust in direction across models and paraphrases. Magnitude was judge-dependent: Sonnet agreed on direction but nearly halved the effect (+13.1 points), with one-directional disagreement. Blinded clinicians characterized the primary judge as a high-sensitivity (1.00), low-specificity (0.55) screen, not a calibrated rate. The gain carried a model-specific helpfulness cost (correct diagnosis 80.3% to 50.3%): near-free for GPT-5.5, near-total for Gemini (-58 points). Matched scaffold controls showed genuine behavior change, not judge circularity. Conclusions: LLM-judged clinical safety effects should be reported as directional and relative, anchored to human review and evaluated jointly with helpfulness, not as calibrated absolute rates. This does not establish clinical deployment readiness.

临床LLM中证据充分性提示的安全增益依赖于评判模型 · AI 热点