离线策略训练:为决策智能体学习临床试验规划

Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents

精选理由

用真实试验数据训练AI规划临床试验,奖励加权行为克隆F1达46.2%,比工具智能体高很多,做医药AI的可以看。

AI 摘要

这项研究将肿瘤药物开发建模为离线决策问题,预测决策日后六个月的试验组合。研究者整合31.7k条公共记录,构成881个离线决策片段,覆盖45个历史项目。对比四种离线目标和四个前沿LLM智能体后,奖励加权行为克隆表现最佳,指示F1达46.2%,严格F1为14.2%;表现最好的工具智能体分别只有25.0%和2.1%。在2025年8月后的数据上,离线训练模型也明显超过未微调基线。

原文 · arXiv cs.LG

Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents

Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. We study this setting by framing oncology clinical development as an offline decision-making problem in which an agent predicts the next six-month trial portfolio of an oncology drug program from information available at the decision date. To support this, we construct a temporal dataset that combines 31.7k heterogeneous public data records, including trial registries, regulatory reviews, sponsor filings, utilization data, and epidemiology, into 881 offline decision episodes across 45 historical programs. We compare four offline objectives: behavioral cloning, reward-weighted behavioral cloning, learned-reward training, and value-based implicit Q-learning against four frontier LLM agents that share a common date-gated retrieval scaffold across held-out drug, sponsor, drug-class, and temporal splits. Models trained offline outperform the non-fine-tuned baselines, particularly in the post-August 2025 contamination-clean holdout. Reward-weighted behavioral cloning performs the best, obtaining 46.2% indication F1 and 14.2% strict F1 against 25.0% and 2.1%, respectively, for the best-performing tool agent on each metric. These results suggest that structured offline learning can teach agents to plan clinical experiments.