GeneICL:面向批量转录组的表格基础模型
GeneICL: A Tabular Foundation Model for Bulk Transcriptomics
用实测表达谱做预训练的小模型,4.2M 参数在 80 个临床任务上赢了大模型,笔记本 CPU 就能跑,做生信的可以看看。
arXiv 论文提出 GeneICL,一个仅 4.2M 参数的表格基础模型,用于批量转录组数据的临床结局预测。它用实测表达谱构建半合成预训练先验,并采用参数高效的循环架构。模型通过基于 Cox 偏似然残差的无训练归约支持右删失生存预测。在 80 个覆盖分类、回归和生存分析的临床预测任务上,GeneICL 在所有基础模型和调优基线中取得最佳综合排名,参数量最多减少 387 倍,推理无需梯度更新,笔记本 CPU 上数秒即可出结果。
GeneICL: A Tabular Foundation Model for Bulk Transcriptomics
Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature correlations, and limited labeled data. Large self-supervised transcriptomic foundation models often fail to outperform simple supervised baselines. Tabular foundation models offer an alternative through in-context learning, but are typically pretrained on generic synthetic data rather than transcriptomic structure. We ask whether transcriptomics-aware pretraining, rather than scale, is the missing ingredient. Towards this end, we introduce GeneICL, a 4.2M-parameter tabular foundation model combining a semi-synthetic pretraining prior built from measured bulk expression profiles with a parameter-efficient recurrent architecture. We further enable right-censored survival prediction via a training-free reduction to regression using Cox partial-likelihood residuals. We evaluate GeneICL on 80 clinical outcome-prediction tasks spanning classification, regression, and survival. Tabular foundation models consistently outperform self-supervised transcriptomic models, while GeneICL achieves the best overall rank among evaluated foundation models and tuned baselines. GeneICL does so with up to 387$\times$ fewer parameters, no gradient updates at inference, and predictions within seconds on a laptop CPU.