论文

HazardWeaver:面向灾害分析智能体的科学路线选择框架

HazardWeaver: Scientific Route Selection for Hazard Analysis Agents

精选理由

一篇把 LLM 智能体用在自然灾害分析上的论文,亮点是让智能体自己判断该用哪种科学方法,还配套了 141 个实例的基准,代码开源可以直接跑。

arXiv 论文 HazardWeaver 把灾害分析中的方法选择问题形式化为状态依赖的科学路线选择。系统包含三个组件:Hazard Knowledge Compiler 提取科学方法适用性的证据关联条件,Hazard Capability Graph 表示可执行的科学能力并检查输入输出兼容性,Hazard Weaver Agent 据此选择路线、执行工作流并随分析状态变化修正决策。配套的 Hazard Weaver Benchmark 含 141 个实例,覆盖 7 个单一灾害领域和 4 类多灾害交互,实验显示 HazardWeaver 优于现有智能体系统,在多条可用科学路线的任务上提升最大。代码已在 GitHub 开源。

原文 · arXiv cs.AI

HazardWeaver: Scientific Route Selection for Hazard Analysis Agents

Understanding and assessing natural hazards is essential for disaster preparedness and risk reduction. Recent advances in large language models have spurred growing interest in AI agents for hazard analysis, particularly their ability to integrate scientific data, models, and tools into automated workflows. However, effective automation requires agents to determine which scientific methods are appropriate for a given event and executable with the available data and tools. As new evidence and execution results become available, these conditions can change, requiring agents to reconsider their choices. We formulate this problem as state-dependent scientific route selection and introduce HazardWeaver. Specifically, HazardWeaver first leverages the Hazard Knowledge Compiler to extract evidence-linked conditions governing scientific applicability, then its Hazard Capability Graph represents executable scientific capabilities and checks compatibility between their inputs and outputs. Using these complementary representations, the Hazard Weaver Agent component selects applicable and executable routes, carries out their workflows, and revises its decisions as the analysis state changes. To evaluate both the scientific outputs and the decisions that produce them, we introduce the Hazard Weaver Benchmark, comprising 141 instances across seven single-hazard domains and four multi-hazard interaction classes. The benchmark accommodates multiple valid scientific routes and evaluates output correctness, route validity, and justified abstention. Extensive experiments on this benchmark show that HazardWeaver outperforms existing agent systems, with the largest gains on tasks with multiple eligible scientific routes. Our code is publicly available at https://github.com/LabRAI/HazardWeaver.