基因调控网络推理的深度与概率模型:PMF-GRN和GLM-Prior

Deep and Probabilistic Models for Gene Regulatory Network Inference

精选理由

想搞基因调控网络但被跨物种问题卡住?这论文用序列预测加概率精修,直接解决迁移和不确定性问题,值得做计算生物学的看看。

AI 摘要

基因调控网络(GRN)从全基因组数据重建面临模型选择启发式、评估参考不完整、先验知识跨物种迁移难等挑战。PMF-GRN将GRN推理转化为概率图模型,通过变分推理实现原理性模型选择和不确定性感知的边估计。GLM-Prior微调预训练Nucleotide Transformer,直接从核苷酸序列预测转录因子-靶基因相互作用,并在酵母、小鼠和人类数据上验证泛化能力。两者结合形成双阶段视角:序列先验提供可迁移初始骨架,概率推理以量化不确定性精调调控估计。

原文 · arXiv cs.LG

Deep and Probabilistic Models for Gene Regulatory Network Inference

Gene regulatory networks (GRNs) link transcription factor (TF) proteins to their target genes, yet reconstructing these networks from genome-wide data remains challenging under practical and methodological constraints. Many methods couple modeling assumptions to a specific inference procedure and rely on heuristic model selection, while evaluation is constrained by incomplete reference networks and point-estimate outputs that lack uncertainty. GRN reconstruction also depends on prior knowledge to constrain TF-gene interactions, yet available priors are often assay-dependent and difficult to transfer across species and less-characterized systems. In this thesis, we develop two complementary frameworks that address these limitations. In the first, PMF-GRN casts GRN inference as a probabilistic graphical model optimized by variational inference, enabling principled model selection and uncertainty-aware edge estimates. In the second, GLM-Prior addresses the prior bottleneck by fine-tuning the pretrained Nucleotide Transformer to predict TF-target gene interactions directly from nucleotide sequence, while generalizing across yeast, mouse, and human settings. Together, PMF-GRN and GLM-Prior motivate a dual-stage view of GRN reconstruction in which sequence-derived priors provide a transferable starting scaffold and probabilistic inference refines regulatory estimates with quantified uncertainty under incomplete evaluation resources.