EpiBench:AI智能体在表观基因组学分析中的可验证基准

EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

精选理由

做基因组学分析的团队终于有了一个可复现的AI能力评估标准——EpiBench揭示了当前最强模型在专业科学判断上的天花板,做生物信息学工具开发或AI+生命科学研究的建议点开看看差距在哪。

AI 摘要

研究人员推出了EpiBench,一个用于短周期表观基因组学分析的可验证基准测试。该基准包含106个评估任务,覆盖CUT&Tag/CUT&RUN、ATAC-seq、ChIP-seq和DNA甲基化等流程。在16个模型-工具组合的5088条有效轨迹中,没有系统通过大部分尝试:GPT-5.5/Pi以45.0%的通过率领先,GPT-5.5/OpenAI Codex以39.9%紧随其后。性能因检测类型而异,许多失败运行仍包含部分正确答案,但任务需要更深入的、检测特定的科学判断时,智能体往往失败。这表明当前AI在需要专业领域知识的复杂分析中仍有明显短板。

原文 · arXiv cs.AI

EpiBench: Verifiable Evaluation of AI Agents on Epigenomics Analysis

We introduce EpiBench, a verifiable benchmark for short-horizon epigenomics analysis. EpiBench evaluates whether agents can make well-defined analysis decisions from realistic workflow states and return deterministically gradable answers. The benchmark includes 106 evaluations across CUT\&Tag/CUT\&RUN, ATAC-seq, ChIP-seq, and DNA methylation workflows. Across 5,088 valid trajectories from 16 model-harness pairs, no system passed a majority of attempts: GPT-5.5 / Pi led at 45.0\% (143/318 attempts; 95\% confidence interval (CI), 36.3--53.7), followed by GPT-5.5 / OpenAI Codex at 39.9\% (127/318 attempts; 95\% CI, 31.6--48.3). Claude Opus 4.8 Max / Pi and GPT-5.4 / Pi each passed 39.0\% (124/318 attempts; 95\% CI, 30.2--47.8 and 31.0--47.0, respectively). Performance varies across assay types, and many failed runs still contain parts of the correct answer. Agents often found the right files and computed useful intermediate results, but failed when the task required deeper, assay-specific scientific judgment.