TrialAtlas 多智能体系统用于临床试验设计与优化
TrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization
朋友,TrialAtlas 这个多智能体系统挺有意思的,专门帮药企做临床试验设计,能从文献和监管历史里学习经验,比 OpenAI 和 Gemini 的效果都好,值得看看。
TrialAtlas 是一个记忆增强的多智能体研究组织,用于临床试验开发规划(CDP)。它通过协调专门负责文献综述、竞争性试验情报、监管先例分析和综合推理的智能体来模拟协作过程。系统从历史临床试验和监管结果(包括新药申请)中学习,以基于积累的经验做出决策。在由 291 份 FDA 完整回复信件构成的基准测试中,TrialAtlas 在缺陷检测任务上达到 50.0% 的 F1 分数,超越最强基线 6.1 分。在技术及监管成功预测任务上,其平衡准确率达 85.3%,F1 分数 84.7%,分别比最佳基线高出 6.7 分和 12.0 分。专家评估显示,TrialAtlas 生成的 86.4% 的担忧被判定为有效,而 OpenAI DeepResearch 为 83.1%,Gemini DeepResearch 为 59.3%。
TrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization
Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely on clinical development planning (CDP) and probability of technical and regulatory success assessment to anticipate development risks, yet these decisions remain labor-intensive and subjective, requiring experts across clinical science, statistics, regulatory affairs, and competitive intelligence to jointly acquire, synthesize, and reason over heterogeneous evidence. Here, we introduce TrialAtlas, a memory-augmented multi-agent research organization for CDP that mirrors this collaborative process by coordinating specialized agents for literature synthesis, competitive trial intelligence, regulatory precedent analysis, and integrated reasoning over trial design and development risk. TrialAtlas further learns from historical clinical trials and regulatory outcomes, including prior New Drug Applications (NDAs), to ground its decisions in accumulated development experience. To evaluate these capabilities in an authentic regulatory setting, we introduce TrialAtlasBench, constructed from 291 FDA Complete Response Letters and spanning three practical tasks: detecting trial design deficiencies, recommending actionable design improvements, and predicting technical and regulatory success. TrialAtlas achieves an F1 score of 50.0% for deficiency detection, outperforming the strongest baseline by 6.1 points, and reaches 85.3% balanced accuracy and 84.7% F1 for prediction of technical and regulatory success, improving over the best baselines by 6.7 points in balanced accuracy and 12.0 points in Cohen's kappa. In expert evaluation, 86.4% of TrialAtlas-generated concerns were judged valid, compared with 83.1% for OpenAI DeepResearch and 59.3% for Gemini DeepResearch.