Hugging Face wants to do for gene discovery what AlphaFold did for protein structures. On October 8, its biology team released Carbon-A, a 1.2-billion-parameter open model (MIT-licensed, 4.6GB of weights) that predicts where protein-coding genes sit in raw DNA sequence — no reference annotation required — alongside the Carbon Annotation Database, or CADB, built by running the model across public genome assemblies at a scale no curated database currently matches.
The model itself is deliberately narrow. It is a eukaryotic DNA model that scores coding probability on both strands within a 98,304-base-pair context window; bacteria and viruses are out of scope, and the team made a point of what it does not do: Carbon-A predicts coordinates, not gene function, and it does not design new DNA. Each prediction ships linked to its source assembly, genomic coordinates, a reconstructed coding sequence, the translated protein and a confidence score, so researchers can filter by precision rather than trust the model blindly.
The database is where the scale lands. The current CADB release covers 48,167 assemblies from 22,617 taxa — roughly 27 trillion base pairs — containing 566.34 million predicted protein-coding loci. Compared with the RefSeq-derived training corpus, that is about 11 times more taxa and 9 times more sequence; against RefSeq's total gene annotations, the candidate count is roughly 16 times larger. Hugging Face says it has annotated about half of its target GenBank set, plans another batch in three weeks, and will make the annotations available through EMBL as well.
The accuracy figures are the company's own and have not been independently verified. Across 42 benchmark genomes — 28 from training snapshots, 14 held out by date — Hugging Face reports a macro-averaged nucleotide F1 of 0.944. At the gene level, the technical report gives precision of 0.792, recall of 0.691 and an F1 of 0.736 with the gene-confidence filter set to 0.10, and the confidence score itself reaches an AUROC of 0.876 for separating exact coding-sequence matches from other predictions. These are reference-agreement numbers, not proof that predicted genes are real.
That proof is arriving in the lab, cautiously. Working with ActiveSite and UCSD, the team read real RNA molecules from cat, Syrian hamster, chicken and Arabidopsis cells and found transcriptional evidence for 239 genes that are missing from RefSeq's reference annotations of those well-studied species. Hugging Face is explicit about the limit: those experiments "do not yet establish that those RNAs are translated into proteins — or what those proteins do."
The honest caveats are worth stating, because the headline number invites over-reading. The 566 million figures are predicted loci — candidate annotations, not 566 million confirmed new genes. The model cannot resolve alternative isoforms, producing a single consensus path per locus. Accuracy on poorly studied species is undocumented, full-length CPU inference is slow, and without fused attention kernels the model card warns of more than 32GB of additional GPU memory. A CADB Explorer is live for researchers to check whether their organism is already covered.
What changes with this release is the reference frame, not the biology. Sequencing has outrun annotation for years — Hugging Face's own framing is that "we can now generate genome assemblies far faster than we can understand them" — and CADB turns that imbalance from a storage problem into a searchable, filterable hypothesis generator. EMBL-EBI's Fergal Martin, whose institute will carry the annotations, called the work a demonstration of "the different opportunities for applying AI to genome annotation." When candidate genes stop being scarce, the bottleneck shifts back to where it has always been: wet-lab validation, one experiment at a time.
Comments (0)
Log in to join the discussion
Log InNo comments yet