New AI Model for DNA Learns from Evolution to Unlock Secrets of the Human Genome

Over two decades have passed since the international scientific community celebrated the milestone achievement of sequencing the first complete human genome. Spanning approximately 3 billion base pairs—the fundamental chemical letters of DNA—this foundational blueprint provided humanity with an unprecedented view of its own biological architecture. Yet, despite the magnitude of that milestone, the precise functional meaning encoded within vast stretches of our genetic code remains profoundly enigmatic.

While historical estimates indicate that only 1 to 2 percent of human DNA actively codes for proteins, the remaining territory has long been casually categorized as "junk DNA." Far from being evolutionary debris, contemporary molecular biology recognizes that much of this non-coding landscape houses vital regulatory elements. These sequences dictate when, where, and with what intensity specific genes are expressed. Decoding these complex regulatory networks represents one of the final frontiers in medical science, holding the potential to illuminate the origins of complex inherited conditions, including cancer, cardiovascular disorders, and neurodevelopmental conditions like autism.

To bridge this knowledge gap, a team of researchers at the University of California, Berkeley, has unveiled GPN-Star, a groundbreaking genomic language model designed to identify high-impact genetic variants. Published on September 9 in the prestigious journal Nature, this development marks a significant departure from conventional artificial intelligence architectures in bioinformatics, offering unmatched computational efficiency alongside superior predictive performance.

Rethinking Genomic AI Through Evolutionary Biology

Genomic language models function conceptually in a manner akin to advanced text-based chatbots, but rather than analyzing natural human languages, they ingest massive corpuses of nucleotide sequences. By evaluating billions of base pairs, these deep-learning architectures leverage pattern recognition to decode the underlying "grammar" of the genome.

However, prior state-of-the-art models—such as the massive Evo 2 framework introduced earlier this year—rely on training protocols that consume staggering amounts of computational infrastructure. Evo 2, capable of generating novel genomic sequences from scratch, required training datasets spanning more than 100,000 distinct species across all domains of life, utilizing 2,000 advanced NVIDIA processors over the course of several months. Such massive computational demands create steep financial and logistical barriers, limiting accessibility for smaller academic laboratories.

GPN-Star fundamentally streamlines this paradigm by changing how input data is curated. Rather than ingesting unaligned genomes from disparate organisms, the UC Berkeley team engineered the model using whole-genome alignments (WGAs). These specialized computational pipelines anchor the genomes of numerous species to a single reference genome, cleanly highlighting conserved sequences versus regions that have diverged rapidly over evolutionary time.

By utilizing human-anchored WGAs alongside comparative datasets for model organisms—including mice, fruit flies, chickens, roundworms (C. elegans), and the model plant Arabidopsis thaliana—the researchers pre-filtered the training data to emphasize functional elements. According to study senior author Yun Song, a professor of computer science and statistics at UC Berkeley and an investigator at the Innovative Genomics Institute, this methodological pivot bypasses the computational noise typically introduced by non-functional junk DNA. Consequently, GPN-Star can be trained in mere days or even hours using only a modest cluster of processors.

Chronology of the Breakthrough

The development of GPN-Star follows a multi-year trajectory of rapid acceleration in computational biology, artificial intelligence, and genomics:

Decoding the ‘grammar’ of the genome with new AI model
  • 2003: The Human Genome Project officially concludes, delivering the first nearly complete sequencing of human DNA, though functional annotation of non-coding regions remains largely incomplete.
  • Early 2020s: Early genomic language models emerge, adopting transformer-based architectures from natural language processing to predict the functional consequences of mutations.
  • Early 2026: Massive cross-species models, such as Evo 2, demonstrate generative genomic capabilities but underscore the heavy computational burden associated with unaligned multi-species datasets.
  • September 9, 2026: UC Berkeley researchers officially publish the GPN-Star framework in Nature, demonstrating that targeted evolutionary data curation can yield superior predictive accuracy for disease variants at a fraction of traditional computational costs.

Unlocking Evolutionary Timescales for Precise Medical Insights

A critical finding of the UC Berkeley study is the correlation between evolutionary timescales and the functional specificity of genomic predictions. Study co-first author Chengzhong Ye, a graduate student in statistics at UC Berkeley, noted that training models across different evolutionary windows optimizes them for distinct categories of genetic variation.

The researchers evaluated three distinct human-anchored WGA scales: one spanning other primates, one covering mammals, and a broader vertebrate dataset. The analysis revealed that models trained on deeper evolutionary timescales—such as vertebrate comparisons—excelled at predicting the pathogenicity of rare genetic variants within protein-coding regions. Because proteins are fundamental to cellular machinery across diverse species, they evolve under stringent constraints over vast spans of time.

Conversely, models trained on tighter, primate-specific evolutionary timescales proved substantially more effective at evaluating genetic variants associated with complex, multi-genic human traits, such as schizophrenia. These conditions are frequently influenced by thousands of distinct mutations scattered across non-coding regulatory regions that have evolved relatively recently in the primate lineage.

"For complex traits, we were surprised and pleased to see that training a model that’s specific to primate genomes—which are more relevant to recent human evolution—really helped us make better predictions," noted Professor Song, who also serves as director of the Berkeley Center for Computational Biology and co-director of the UC Berkeley-UCSF Bakar Computational Biomedicine Initiative.

Empirical Validation and Broad Scientific Implications

Beyond publishing the underlying neural network architecture, the UC Berkeley research team has released comprehensive, genome-wide predictions generated by GPN-Star. These public annotations highlight specific genetic variants projected to exert the strongest influence on human traits and inherited disorders.

In modern biomedical research, high-throughput functional assays allow scientists to test the biochemical effects of genetic alterations. However, experimental evaluation of all 3 billion base pairs in the human genome remains physically impossible due to temporal and financial constraints. By providing a curated roadmap of high-probability pathogenic variants, GPN-Star provides the scientific community with a powerful prioritization tool.

"People have developed really creative tools for assaying the impact of genetic variants, but they cannot experimentally test every single variant in the genome," Professor Song emphasized. "We believe our predictions will help to prioritize the experiments that could have the greatest impact on human health."

This emphasis on accessibility is intentional. Because GPN-Star requires minimal computational overhead to train and modify, the research team anticipates that global laboratories will readily adapt, refine, and build upon the framework. Study co-first author Gonzalo Benegas underscored this collaborative philosophy, stating that broader engagement from the global research community will iteratively strengthen the model’s capabilities over time.

Funded in part by the National Institutes of Health, the GPN-Star project represents a maturing intersection between evolutionary biology and artificial intelligence. By allowing natural selection to perform the heavy lifting of biological filtering, UC Berkeley’s approach demonstrates that computational efficiency in genomics does not require sacrificing predictive precision. As biomedical researchers begin integrating these genome-wide annotations into clinical and academic workflows, GPN-Star stands poised to accelerate the translation of raw genetic data into tangible diagnostic and therapeutic breakthroughs.