AI Is Learning to Predict RNA Structure, Opening New Doors for Gene Therapy and CRISPR
A team at UC Berkeley's Innovative Genomics Institute has made significant progress using artificial intelligence to predict how RNA molecules fold into three-dimensional structures, a breakthrough that could accelerate the development of gene therapies, CRISPR tools, and other biomedical applications. The research, published this week in the journal RNA, addresses a long-standing challenge in molecular biology: while AI has revolutionized protein structure prediction, RNA has remained largely out of reach due to limited training data and the molecule's inherent flexibility.
Why Has RNA Structure Prediction Lagged So Far Behind Protein Prediction?
The success of AlphaFold2, which earned its creators the Nobel Prize in Chemistry in 2024, demonstrated that deep learning models could predict protein structures with remarkable accuracy. However, RNA presented a fundamentally different problem. Proteins benefited from massive, publicly available datasets of sequences and solved structures that researchers could use to train AI models. RNA, by contrast, has far fewer high-quality structural datasets available.
Beyond data scarcity, RNA molecules themselves are far more challenging to model computationally. Unlike DNA's rigid double-stranded structure, RNA consists of a single strand that can fold into complex, flexible shapes. A single building block in an RNA molecule is roughly three times larger than a protein building block, giving RNA substantially more freedom to bend and twist. To complicate matters further, individual RNA molecules can adopt multiple stable structures, each potentially performing different functions.
"RNA has always lagged a couple decades behind protein biology because there's fewer people working on it, and it's not as directly translational into bedside work," said Conner Langeberg, a postdoctoral researcher in the Doudna and Cate labs at the Innovative Genomics Institute. "But there's a lot of use in knowing how RNAs work. They are some of the most ancient machines in the cell. Ribosomes, which make the proteins in our cells, are RNA-dependent machines. CRISPR is another great example: without RNA, these types of tools wouldn't be accessible to us."
Conner Langeberg, Postdoctoral Researcher, Innovative Genomics Institute at UC Berkeley
How Did Researchers Overcome the Data Shortage?
The breakthrough came from an unexpected source: a 2024 hackathon organized by the Innovative Genomics Institute. During that event, Langeberg worked on curating high-quality RNA structure datasets to train machine learning models. While doing so, he realized that vast amounts of untapped structural data for RNA molecules existed scattered across public databases, simply waiting to be organized.
Langeberg then undertook an ambitious data-gathering effort, using bioinformatic pipelines to search through over 4,000 different families of RNA, including genomes from bacteria, archaea, viruses, and eukaryotes. This systematic search produced a new large-scale dataset called RNASSTR, which the team made freely available to the scientific community.
The team's key innovation was how they used this dataset to retrain existing machine learning algorithms. Rather than simply feeding all the data into models, they kept related RNA structures together during training, preventing the models from "cheating" by memorizing similar sequences. They then tested the retrained models against the original versions using data that had been set aside for validation. The results showed that the retrained models performed significantly better on specific classes of RNA structures.
Steps to Advance RNA Structure Prediction
- Data Curation: Systematically search public databases to identify and organize high-quality RNA structural data across diverse organisms and RNA families, creating unified datasets that researchers can access freely.
- Model Retraining: Retrain existing machine learning algorithms using expanded datasets while maintaining the relationships between related structures to prevent models from learning shortcuts rather than genuine patterns.
- Computational Optimization: Develop new model architectures designed to handle increasingly large datasets more efficiently, reducing the computational expense of training on robust, comprehensive data.
- Validation and Testing: Rigorously test retrained models against held-out datasets to identify which RNA classes show the most improvement and where further refinement is needed.
What's Next for RNA AI Models?
The current work represents an important stepping stone, but the ultimate goal remains creating an AlphaFold-like tool that can generalize across all RNA structures. One major hurdle the team encountered was the sheer computational expense of training models on increasingly large datasets. To address this challenge, researchers are now developing a new architecture called "Lyra-TransPred," designed to accelerate training on larger, more robust datasets with the hope that it will allow models to generalize across more classes of RNA structures.
"It's an exciting step forward. Hopefully from there we can make progress towards generalizable structure predictors of RNA," noted Langeberg.
Conner Langeberg, Postdoctoral Researcher, Innovative Genomics Institute at UC Berkeley
The practical implications are substantial. Accurately predicting RNA structure from sequence alone would enable researchers to design new RNA-based tools for gene therapy, improve CRISPR systems, and engineer RNA molecules for therapeutic purposes. The RNASSTR dataset is already available online, allowing scientists worldwide to build on this foundation.
Meanwhile, recognition of the importance of genomic research continues to grow. Dan Landau, a researcher at the New York Genome Center, recently received the 2026 European Society for Medical Oncology Award for Translational Research, recognizing his pioneering work in cancer genomics and computational approaches that translate genetic discoveries into clinical applications. His work on somatic evolution, single-cell multi-omics, and circulating tumor DNA detection demonstrates how genomic and computational advances are increasingly moving from the laboratory into patient care.
The convergence of improved AI models, expanded datasets, and growing clinical translation suggests that RNA structure prediction could soon become as routine and powerful a tool as protein structure prediction has become, potentially unlocking new therapeutic possibilities for diseases that have long resisted treatment.
" }