产品精选

Thom Wolf 团队开源 Carbon-A 模型,直接从 DNA 中定位基因

精选理由

EleutherAI 联创放出开源基因定位模型 Carbon-A,扫了 2 万多个物种找到 5 亿多个候选基因,还附数据库,搞生信的可以看看。

Carbon-A 是一个开源模型,直接在 DNA 序列中找出基因位置,无需依赖与小鼠、果蝇等模式物种的比对。团队用它构建了覆盖 22,617 个物种、共 5.66 亿个候选基因的 Carbon Annotation Database。与 Active Site 和 UCSD 合作,他们在猫、鸡、仓鼠和拟南芥等常见物种的参考注释中找到了 239 个此前缺失的基因 RNA 证据。模型不做 DNA 设计、不预测基因功能,只标注基因位置。

原文 · Thomas Wolf

Scientists have sequenced the genomes of thousands of species. But for most of them, nobody knows *where the genes are* Today we're releasing Carbon-A, an open model that finds genes directly in DNA. We used it to create a database of 566 million candidate genes across 22,617 species, from fungi to mammals (that we are sharing as well) Why bother? A few reasons 1. Elephants rarely get cancer. Part of the explanation turned up in their genes: extra copies of a tumor-suppressor gene humans have only one of. You can't ask that kind of question about a species until someone, or some tool, has found its genes. 2. The first GLP-1 drug was based on a peptide from Gila monster venom. Nobody would have put the Gila monster on a priority list. Same for wild relatives of crops, which carry resistance to drought and disease. Many of these species have never been annotated. 3. Most of what we know about genes comes from about a dozen model species, like mice, flies and yeast. That focus worked remarkably well, but annotation pipelines still lean heavily on comparisons with them, which makes genes unique to other species easy to miss, and those can be the most interesting ones. Carbon-A belongs to a newer family of tools that read DNA directly. It learned from known genomes, but it doesn't need a close relative to read a new one. 4. Even familiar genomes still have gaps. With our partners at Active Site and UCSD, we found RNA evidence for 239 genes missing from the reference annotations of species as common as cats, chickens, hamsters and a lab plant. 5. AlphaFold can predict the shape of a protein, but only once someone has found the gene that makes it. In a species with no annotated genes, it has nothing to work with. Some notes on safety: Carbon-A doesn't design DNA or predict what a gene does. It marks where genes are in DNA. This is one of the reasons we think releasing it openly is the right call. More annotations also help health research: many species that carry or cause disease are among those covered, and work on controlling the diseases they spread often starts from their genes. Model and database and more details in Georgia's thread 👇 Georgia Channing @cgeorgiaw Today, we're releasing the Carbon Annotation Database and the model behind it, Carbon-A. We've used this model to discover 566.34 million new gene candidates across 22,617 species, thousands of which have never been studied like this before. We wet-lab validated a number of these new genes in well-studied species, such as in cats, chicken, and arabidopsis (plant), with our partners at Active Site and @UCSD , with more extensive wet lab results on previously-unstudied genomes coming soon. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 6 ❤️ 36 👀 1995 📊 6 ⚡