做生物信息学或组学数据降维的团队,终于有了一个标准化的 HPO 测试场——BBOmix 帮你省去从头调参的试错成本,做自编码器研究的可以直接用它验证方法。
BBOmix 是首个针对真实生物数据的无监督表示学习超参数优化(HPO)开源表格基准。它包含来自 TCGA 和 SCHC 数据集的 105,000 次评估,涵盖四种自编码器架构和七种多组学模态。该基准量化了重建损失与下游任务性能之间的相关性,并评估了多种 HPO 方法,为无监督生物表示学习研究建立了严格基线。
BBOmix: A Tabular Benchmark for Hyperparameter Optimization of Unsupervised Biological Representation Learning
The rapid advancement of high-throughput sequencing has led to large, high-dimensional omics datasets. Deep unsupervised learning architectures, particularly Autoencoders (AEs), are increasingly used for dimensionality reduction and representation learning in this domain. However, AEs are highly sensitive to architectural choices and hyperparameters, and unsupervised optimization typically relies on reconstruction loss, which may be a poor proxy for downstream utility. Exhaustive hyperparameter optimization (HPO) is computationally expensive, leading researchers to frequently rely on suboptimal default configurations. To democratize access to large-scale unsupervised HPO research, we introduce $\textbf{BBOmix}$, the first open-source tabular benchmark for unsupervised representation learning on real-world biological data. Our benchmark includes 105,000 evaluations across four AE architectures and seven multi-omics modalities from the TCGA and SCHC datasets. We quantify the correlation between reconstruction loss and downstream task performance and provide an extensive evaluation of state-of-the-art single-fidelity, multi-fidelity, and transfer learning HPO methods, establishing a rigorous baseline for future research in unsupervised biological representation learning.