BRIE:可持续更新的电子病历信息检索基准
A Living Benchmark for Information Retrieval from Electronic Health Records
19 位医生参与验证的 EHR 检索基准 BRIE,能自动生成题目并持续更新防泄漏,测了 9 个模型,做医疗 AI 评测的可以看看。
arXiv 论文提出从纵向电子病历(EHR)笔记自动生成问答对的可扩展框架,构建出可长期维护的基准 BRIE。19 名临床医生对生成器进行了验证。该基准在 9 个 LLM 和 5 种推理策略上做了评测,发现最先进的系统经常遗漏临床重要信息,在需要跨多份文档和多次就诊做综合的题目上尤为明显。由于生成器本身经过验证,BRIE 可持续刷新题目内容以防评测泄漏,并生成反映医生推理差异的多个参考答案。
A Living Benchmark for Information Retrieval from Electronic Health Records
Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes. Nineteen clinicians validate the benchmark generator, producing the Benchmark for Retrieving Information in EHRs (BRIE), a continuously maintainable evaluation dataset. Across nine LLMs and five inference strategies, state-of-the-art systems frequently omit clinically important information, particularly for questions requiring synthesis across multiple documents and encounters. Because the generator itself is validated, BRIE supports evaluations that static benchmarks cannot, including the generation of multiple answers that reflect variation in clinician reasoning for robust performance assessment and continuously refreshing benchmark content to guard against leakage. Our results demonstrate that scalable benchmark generation enables rigorous, up-to-date evaluation of clinical LLMs as they are deployed in rapidly evolving healthcare settings.