Hugging Face 把 DNA 分析从黑盒 API 拉到了本地,做生物信息学或个性化健康研究的开发者可以直接在笔记本上跑基因组模型,值得试试。
Hugging Face 发布了名为 Carbon 的开源 DNA 基础模型,包含开放权重、训练代码和数据管道。该模型专为下游生物学任务设计,可微调或持续预训练。Carbon 比同尺寸最佳模型快 275 倍,能在单 GPU 上不到 2 天处理整个人类基因组,甚至可在笔记本电脑上本地运行。其核心技术是 DNA 原生分词器,将序列分割为 6 碱基块以提升效率,同时保留单碱基分辨率。此举旨在推动生物学 AI 的透明化和本地化,避免个人健康数据依赖黑盒 API。
The future of biology shouldn’t stay behind black-box APIs. Especially when it touches personal heal...
The future of biology shouldn’t stay behind black-box APIs. Especially when it touches personal health. Whether you’re @bryan_johnson measuring every biomarker, or @sytses openly sharing and analyzing his own immune-genetics data, you need open, local, transparent AI. @huggingface wasn’t created to be a biology company. It’s not the most obvious focus for us. But it feels too important not to do something. That’s why we built and released Carbon 🧬: a frontier DNA base model with open weights, training code and data pipeline, designed to be fine-tuned or continually pretrained for downstream biological tasks. Carbon is 275x faster than the next best model at its size. Fast enough to run locally on your laptop. Powerful enough to process a whole human genome on a single GPU in less than 2 days. The technical unlock: a DNA-native tokenizer that splits sequences into 6-base chunks for efficiency, while preserving single-base resolution during training and inference. More people able to inspect, run, fine-tune, improve and build on top of the models shaping biology. Open weights: huggingface.co/collections/Hu… q Dataset: huggingface.co/datasets/Huggi… P Demo: huggingface.co/spaces/Hugging… b Let's go open AI biology! 💬 9 🔄 14 ❤️ 62 👀 3333 📊 19 ⚡