三种AI医疗文书系统审计结果

One note in three: a verified census of three deployed AI scribes, and the instrument that counted it

精选理由

三种AI医疗文书系统被审计,近三分之一存在错误,尤其在过敏信息和患者身份方面。

AI 摘要

研究人员对三种商业AI医疗文书系统进行了审计,分析了565份临床记录。AI系统在31.3%的笔记中存在验证失败,主要集中在过敏和药物信息、患者身份虚构以及电话咨询中历史记录错误。在没有患者记录的情况下,错误率为24.8%。不同审核标准和模型家族会导致错误率在28%到97%之间波动。

原文 · arXiv cs.AI

One note in three: a verified census of three deployed AI scribes, and the instrument that counted it

Ambient AI scribes draft clinical notes under the reassurance that a clinician signs every note. We audited three commercial AI scribes on the same 142 consultations: 565 notes from recorded UK primary-care and US ambulatory encounters plus authored scenarios. Twelve discovery passes proposed 13,678 candidate errors; the 5,898 clearing an importance filter went to an adversarial panel of two models from different families, each told to refute what it could, and 618 survived. One note in three (31.3% [27.0, 35.6]) carries a verified failure, concentrated in allergy and medication information, invented patient identity, and history written up as examination on telephone consultations that can contain none. No product was given a patient record; setting aside the two classes a record would have prefilled, invented identity and dates, the rate is 24.8% [20.8, 29.0]. One failure mode did not fit our scheme, drawn from published scribe-error taxonomies: a treatment the clinician retracts, recorded as delivered care. Two clinicians adjudicated blind, disjoint samples: a physician author upheld 20 of 21 findings (95.2% [77.3, 99.2]) and an independent clinician, not an author, 12 of 12 ([75.8, 100]); both judged every sampled refusal genuine. A failure rate depends on the instrument as much as the scribes. With model, evidence and settings fixed, the review instruction alone moves the share of candidates verified from 9.3% to 79.0%, and the reviewing family moves it too: alone at that instruction the gentler flags 54.8% of notes against 27.8%. Between 28% and 97% of sampled notes carry a failure depending on the standard. Published audits disagree among themselves by a margin instrument differences alone can produce: omission is 54-86% of their errors against our 23.1%. We release all 618 findings with transcript-side evidence, every prompt and model version, and the re-runnable pipeline.