EvoSCM 让 AI 能像科学家一样建立、测试和修正因果理论,在物理发现任务中表现优异。
研究人员提出EvoSCM框架,为科学智能体配备显式结构化因果模型。该系统维护竞争性假设种群,通过 abduction 推理潜在机制、设计区分性干预、进行可证伪预测。在 DiscoverPhysics 基准测试中,EvoSCM 持续超越基线模型,提供更准确的解释和预测,同时更有效地利用实验交互。
EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation
Scientific agents must learn not only how to reason, but also what to believe. However, existing LLM agents typically express scientific hypotheses in free-form text, leaving their beliefs implicit and difficult to test or revise. We introduce EvoSCM, which equips scientific agents with explicit structural causal models that evolve as new experimental evidence is collected. EvoSCM maintains a population of competing SCM hypotheses, each encoding a candidate causal explanation of the environment, and evolves them through a closed discovery loop. In each round, the agent abduces latent mechanisms from accumulated evidence, designs discriminative interventions, and commits to falsifiable predictions that it tests through experimentation. Discrepancies between prediction and observation are inductively distilled into correction rules that revise the causal structures and mechanisms of each hypothesis, and the agent then deductively validates the revised population against accumulated evidence and structural consistency to guide the next round. We evaluate EvoSCM on DiscoverPhysics, a benchmark requiring agents to uncover the hidden dynamics of noncanonical physical worlds through experimentation. EvoSCM consistently improves scientific discovery over baselines, yielding more accurate explanations and predictions while making more effective use of experimental interactions.