论文精选

前沿LLM智能体突破自然表型本体注释瓶颈

Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

精选理由

做生物信息学或本体工程的研究者终于有了可扩展的自动化方案——LLM智能体直接对标人类专家水平,建议点开看具体实现和评估细节。

AI 摘要

表型注释是将自由文本描述链接到本体术语的关键步骤,但传统上依赖高训练专家,难以规模化。本研究使用Anthropic和OpenAI的五个前沿LLM作为“智能体策展人”,在自包含工作空间中提供原始论文PDF、注释指南和本体文件,评估其与人类策展人的一致性。结果显示,所有智能体均达到原始研究中三位训练人类策展人的一致性范围,最佳智能体接近但未超越最佳人类策展人,且在所有指标上大幅优于传统NLP工具。这表明LLM智能体有潜力自动化表型注释,缓解本体策展瓶颈。

原文 · arXiv: Anthropic

Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

Linking free-text phenotype descriptions to ontology terms, typically referred to as phenotype annotation, is essential for the cross-study integration of comparative morphological data. This labor intensive process has heavily relied on highly trained human experts, which makes it challenging to scale and thus a key bottleneck. Dahdul et al. (2018) established a Gold Standard (GS) of Entity-Quality (EQ) annotations across seven phylogenetic studies and used it to evaluate three human curators and the Semantic CharaParser NLP tool with ontology-based semantic similarity metrics; they reported that machine-human consistency was significantly lower than inter-curator (human-human) consistency. Here we revisit that benchmark with five frontier hosted LLMs from Anthropic and OpenAI, each operating as an "agentic curator" within a self-contained workspace that supplies the source publication PDF, the same annotation guide used by the original human curators, the four project ontologies (UBERON, PATO, BSPO, GO), and a validation script. Evaluated against the same Gold Standard, every agent fell within the range of inter-curator variability of the three trained human biocurators of the original study; the best performing agents approached but did not reach the best performing human curator. Agents substantially outperformed Semantic CharaParser on all four metrics.