从病理报告中提取H. pylori:多智能体系统nMAS的98.61%准确率验证

Finding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy Reports

精选理由

这篇论文展示了一个能精准从病理文本中抓取H. pylori证据的多智能体系统,准确率98.61%,比人工快几十倍,适合医疗信息提取场景。

AI 摘要

新加坡数据显示约31%人口有幽门螺杆菌感染证据。研究者使用Nimblemind多智能体系统(nMAS)处理54份去标识化胃活检病理报告,评估4个临床二元字段(胃活检、活检状态、H. pylori阳性、H. pylori相关胃炎)。在216个特征-案例决策中,nMAS正确分类213个(98.61%准确率),与UMA-style MiniMax M2.5对比表现相似,但nMAS提供统一报告级输出与支持源句。假设人工审查每份5分钟,nMAS每份5秒,1000份报告处理时间可从83.3小时降至1.4小时,节省约6100美元。

原文 · arXiv cs.AI

Finding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy Reports

Data from Singapore indicated that about 31% of the population had evidence of Helicobacter pylori infection. Persistent H. pylori infection is associated with chronic active gastritis and peptic ulcer disease, and its eradication is key to gastric cancer prevention. However, evidence supporting \textit{H. pylori} positivity and H. pylori-associated gastritis may be distributed across heterogeneous coded and free-text report fields and may require contextual interpretation of assertion and negation, limiting keyword search, and making manual review difficult to scale. We conducted a retrospective pilot evaluation of the Nimblemind Multi-Agent System (nMAS), a field-name-driven, evidence-linked extraction workflow, using 54 de-identified gastric biopsy pathology reports from a large healthcare system in Singapore. Four clinician-scoped binary fields were evaluated: gastric/stomach biopsy, biopsy status, H. pylori positivity, and H. pylori-associated gastritis. Across 216 feature-case decisions, nMAS correctly classified 213, corresponding to 98.61% overall accuracy. A separately implemented UMA-style MiniMax M2.5 comparator produced similar aggregate and per-field classification metrics. Although predictive performance was similar, nMAS maintained unified report-level outputs with supporting source sentences; the demonstrated contribution is therefore workflow integration and traceability rather than predictive superiority. Under an illustrative, unmeasured scenario, reviewing 1,000 reports at five minutes per manual review versus five seconds per evidence-linked verification would reduce review time from 83.3 to 1.4 staff-hours, corresponding to 81.9 staff-hours and about USD~6,100 in potential staff-time value. Larger multi-institutional studies should evaluate evidence-span correctness, clinician verification time, and generalizability.