A New Roadmap for Understanding Gene Regulation
by Micaela Harris-Kim
July 20, 2026
Transcriptional enhancers are sections of a DNA strand that determine if a gene will transcribe into a particular RNA. Many genetic changes occur in the transcriptional enhancers rather than the gene themselves. Understanding gene regulation is pivotal in understanding genetic changes that are linked to disease. Furthermore, rise of predictive models in clinical settings presents opportunities for deeper understanding on how exactly transcriptional enhancers affect the genome and influence disease.
In a study recently published in Nature, led by researchers at Stanford University, including first authors Andreas Gschwind, PhD, Kristy Mualim, Maya Sheth and senior authors Jesse Engreitz, PhD, Lars Steinmetz, PhD, William Greenleaf, PhD and Anshul Kundaje, PhD, created a comprehensive database of transcriptional enhancers as part of the ENCODE project. This resource contains more than 92 million enhancer-gene regulatory interactions across 1,458 biosamples, covering 369 cell types and tissues.
Initially, the researchers assembled a benchmarking framework using ground truth enhancer-gene pairs from genetic perturbations, including CRISPR experiments. Using these data to train and evaluate their approach, they developed a prediction model called ENCODE-rE2G that outperforms other existing models. The study shows that machine learning, in combination with large-scale experimental data, can improve prediction of how transcriptional enhancers regulate genes.
The investigators then applied the model to build a reference map of enhancer-gene interactions across the human genome. Revealing how genes are regulated, the map shows that the regulatory landscape of some genes is much more complex than previously thought. The class of the gene's promoter and cooperation between nearby enhancers play important roles in shaping how genes are expressed.
Importantly, this reference map provides a valuable resource for understanding how genetic variants impact disease in different cell types.
This work marks an exciting intersection of artificial intelligence and genomics, furthering understanding in disease mechanisms and clinical applications. In addition, ENCODE’s framework can be adapted to address other questions about gene regulation, such as identifying the genes controlled by regulatory proteins or understanding how genetic variants influence gene activity. Ultimately, more complete maps of gene regulation will provide a valuable foundation in improving disease diagnosis, enhancing clinicians’ ability to manipulate gene expression, and understanding further the biological effects of genetic variation.
Additional Stanford University investigators include Anthony Tan, James Galante, Hank Jones, X. Rosa Ma, David Yao, Dulguun Amgalan, Chad Munger, Suhas Rao, Benjamin Doughty, Kalina Andreeva, Tri Nguyen and Michael Bassik