代谢组学专用大模型MetaboLLM:整合生化知识构建预测性代谢物图谱

MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

精选理由

做代谢组学或临床预测的可以看看,MetaboLLM把生化知识变成图谱,两个预测任务的AUC都挺能打。

AI 摘要

研究团队提出代谢组学专用大模型MetaboLLM,通过持续预训练、监督微调和结构化检索整合生化知识。配套的MetaboLLM-GIN将生成描述转换为代谢物图,用于患者级预测。在四个骨干模型家族上,MetaboLLM在代谢组学知识、关系和描述任务上优于对应基础模型和医学适配模型。MetaboLLM-GIN在冠状动脉搭桥术后应激性高血糖预测中AUC达0.8616,在绝经后激素方案分类中AUC为0.8123,超过传统模型和替代图构建方法。模型解释还产生了具有生物学意义的发现。

原文 · arXiv cs.LG

MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed MetaboLLM, a metabolomics-specialized large language model adapted through continual pretraining, supervised fine-tuning, and structured retrieval, together with MetaboLLM-GIN, which converts generated biochemical descriptions into metabolite graphs for patient-level prediction using a graph isomorphism network. Across four backbone families, MetaboLLM outperformed corresponding base and medically adapted models on metabolomics knowledge, relational, and description tasks, and transferred to an external public benchmark. MetaboLLM-GIN achieved the highest AUC for stress hyperglycemia prediction after coronary artery bypass grafting (0.8616) and postmenopausal hormone-regimen classification (0.8123), outperforming conventional models, alternative graph constructions, and graphs generated from unadapted or non-retrieval LLM configurations. Model interpretation further produced biologically meaningful findings in both applications. These results show that domain-specialized language models can organize heterogeneous biochemical knowledge into predictive and interpretable metabolite graph representations.