单细胞组学研究者终于有了系统评估翻译模型的工具——scTranslation覆盖了数据、指标和影响因素,做多模态分析的团队可以直接用这个基准来对比方法,省去自己搭建评估流程的麻烦。
单细胞多组学数据同时测量多种模态,但实验成本高、噪声大,催生了多种计算翻译方法。然而,现有方法缺乏系统性的基准评估。为此,研究者提出了scTranslation基准,包含多样化数据集、集成最新模型并提供全面评估指标。该基准还评估了特征选择、特征质量和小样本设置等影响因素,这些因素此前很少被系统研究。通过大规模实验,scTranslation揭示了多项重要发现,为未来研究开辟了新方向。基准已开源,代码可在GitHub获取。
scTranslation: A Comprehensive Benchmark for Single-Cell Multi-Omics Modality Translation
Simultaneous measurement of multiple omics modalities in single cells enables researchers to gain a more comprehensive understanding of cellular states and regulatory mechanisms. However, due to high experimental costs, significant noise, and incomplete modality coverage, a variety of computational methods for modality translation have emerged in recent years. Despite the development of translation models, there is still a lack of systematic benchmark evaluation in terms of datasets, evaluation metrics, and influencing factors. To address this, we present scTranslation, a comprehensive benchmark for single-cell multi-omics modality translation tasks. It includes diverse translation datasets, integrates state-of-the-art models, and provides a comprehensive evaluation metrics. In addition, we assess model performance under different scenarios, such as feature selection, feature quality, and few-shot settings. These factors significantly affect model performance but have rarely been systematically studied before. Leveraging this benchmark, we conduct a large-scale study of current methods, report many insightful findings that open up new possibilities for future development. The benchmark is open-sourced to facilitate future research. The code is anonymously released at https://github.com/Bunnybeibei/scTranslation.