Doctorina临床AI系统诊断表现超越医生
Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics
Doctorina在初级护理诊断上全面超越人类医生,比第二名Kimi K3表现更好。
Doctorina在150个合成波兰语初级护理咨询中达到82.0% Top-1诊断一致性,比医生的57.0%高出25个百分点。在149个病例对比中,检查和治疗得分分别为89.4比66.9和83.7比61.2。Kimi K3在诊断中排名第二,Claude Opus 5在管理评估中领先。Doctorina在所有六个组中诊断点估计值最高。
Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics
Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 concordance versus 57.0% for physicians (difference, 25.0 percentage points; 95% confidence interval, 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Across 149 case pairs, normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Doctorina had the highest diagnostic point estimates among all six groups; Kimi K3 ranked next, while Claude Opus 5 led the closely spaced management estimates of Opus, Doctorina and Kimi. A second Doctorina execution reproduced the advantages over physicians across all outcomes. Doctorina's advantage over physicians therefore extended from primary-diagnosis selection to higher-rated diagnostic workup and initial treatment after adaptive consultation.
- Fireworks AI09-08 18:21原文