罕见病诊断难?RareLens利用多个大模型的差异推理,在各阶段都超越GPT-5等顶尖模型,准确率还高。
罕见病影响3.5%至5.9%人口,超70%患者被误诊。RareLens系统通过对齐异构大模型的分歧性推理,在罕见病全病程中提供决策支持,包括初诊风险筛查、诊断、治疗规划和预后。该系统在包含157,525例病例(覆盖33个Orphanet类别、7000+疾病)的RareBench上评估,各阶段均超越GPT-5、DeepSeek-R1、Claude-3.7-Sonnet和Gemini-2.5-Pro。RareLens筛查AUC达0.917,诊断和治疗top-1准确率分别为65.5%和89.8%。外部研究(1287例、23名医生)显示自主RareLens和医生辅助RareLens均显著优于无辅助医生。
RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning
Rare diseases collectively affect an estimated 3.5% to 5.9% of the population, yet more than 70% of patients are misdiagnosed and many endure years of evaluation before a diagnosis is reached, because early presentations are nonspecific and relevant expertise is scarce and unevenly distributed. Artificial intelligence could provide support, but existing systems address isolated stages of care, overwhelmingly diagnosis. They typically depend on the results of downstream investigations, and they treat the variability between models as noise to be eliminated. Here we present RareLens, a system that supports clinical decision-making across the entire rare disease trajectory by exploiting this variability. When heterogeneous large language models evaluate the same case, they generate divergent but complementary reasoning, which RareLens aligns and calibrates into a single convergent, actionable decision at each stage. Four coordinated modules perform primary-visit risk screening, diagnosis, treatment planning and prognosis. Developed and evaluated on RareBench, a real-world dataset of 157,525 cases spanning all 33 Orphanet categories and more than 7,000 conditions, RareLens outperformed every frontier model tested, including GPT-5, DeepSeek-R1, Claude-3.7-Sonnet and Gemini-2.5-Pro, at each stage. It achieved an area under the curve of 0.917 for screening and top-1 accuracies of 65.5% and 89.8% for diagnosis and treatment. In an external study spanning 1,287 cases and 23 physicians, autonomous RareLens and physicians assisted by RareLens both substantially outperformed unaided physicians. These findings indicate that aligning divergent model reasoning, rather than scaling a single model, offers a generalizable strategy for high-uncertainty clinical decision-making.