Why AI's Cancer Breakthroughs Depend on Reading RNA Like a Complete Message
Artificial intelligence is transforming cancer treatment, but researchers have discovered a critical gap: the genetic data AI learns from is often incomplete, like trying to understand an email after reading every third sentence. A new wave of long-read RNA sequencing technology is filling that gap by revealing the complete molecular instructions that determine whether a patient will respond to therapy, giving AI models the high-quality biological information they need to make accurate predictions.
Why Current Cancer Diagnostics Miss the Real Picture?
For decades, cancer biology has relied on DNA sequencing to guide treatment decisions. Doctors test for genetic mutations and protein markers to decide which targeted therapies a patient should receive. But DNA tells only part of the story. It shows what is possible in a cell, not what is actually happening.
Consider HER2-positive breast cancer, one of the most common precision medicine success stories. Patients are tested to see if their tumors produce the HER2 protein, and those who test positive receive drugs like trastuzumab designed to bind to HER2. However, a patient's tumor might produce multiple versions of the HER2 protein, called isoforms, created through a process called alternative splicing. Some of these isoforms contain the drug-binding site; others do not. A standard DNA test cannot distinguish between them.
"The biomarkers currently used in diagnostics typically don't directly measure the drug-binding site or the underlying biology that's required for response," said Richard Kuo, CEO of Wobble Genomics.
Richard Kuo, CEO, Wobble Genomics
This means a patient could receive an expensive, toxic therapy that is designed to work but cannot actually interact with the specific protein variant their cancer is producing. The drug fails not because it is ineffective, but because the diagnostic test missed a critical detail about what the cancer is actually doing.
How Does RNA Reveal What DNA Cannot?
RNA is the intermediate step between DNA and protein. While DNA is the blueprint, RNA is the actual message being read and acted upon inside cells. By sequencing RNA directly, researchers can see which genes are actively being expressed, which protein variants are being produced, and how a cancer is functionally behaving in real time.
Traditional RNA sequencing has relied on short reads, which means researchers reconstruct full transcripts from fragments, much like assembling a puzzle from pieces. Long-read RNA sequencing, by contrast, reads complete RNA molecules directly, revealing the full message without gaps or assumptions.
Wobble Genomics, a spinoff from the University of Edinburgh, has developed a technology called ATLUS that captures circulating long RNA (clRNA) from blood samples. These are full-length RNA molecules released by cancer cells into the bloodstream, often packaged in tiny vesicles that cells use to communicate with their environment. By reading these complete molecules, researchers can identify which specific protein isoforms a tumor is producing and whether those isoforms contain the drug-binding sites required for targeted therapies to work.
In validation testing, the ATLUS platform achieved a limit of detection below 0.05 parts per million, meaning it can identify extremely rare tumor-derived signals from a single blood draw. The platform also demonstrated 100 percent molecular specificity, with no false positive detections of cancer-specific fusion transcripts in the company's validation dataset.
How AI Models Benefit from Complete Biological Data?
Artificial intelligence excels at recognizing patterns, but only when those patterns accurately reflect biology. Sequencing errors, incomplete genome assemblies, fragmented transcripts, and missing isoforms introduce uncertainty into AI models. Rather than learning genuine biological relationships, algorithms may instead learn artifacts of the underlying data.
Modern AI applications in drug discovery increasingly integrate multiple layers of biological information, called a multiomic approach. This means combining genomic data (DNA sequence), transcriptomic data (RNA expression), epigenomic data (chemical modifications to DNA), proteomic data (protein measurements), imaging, and clinical outcomes into a single model.
When researchers provide AI with high-quality, complete data from long-read sequencing, the models can discover more meaningful biological patterns. Instead of compensating for missing or inaccurate information, AI can focus on learning genuine relationships between genetic variation, gene expression, and disease.
Steps to Strengthen AI-Driven Cancer Research
- Integrate Multiple Data Layers: Combine genomic sequence, epigenomic modifications (like DNA methylation), transcriptomic profiles, and clinical outcomes into a single AI model rather than training on isolated data types alone.
- Use Long-Read Sequencing for Complete Transcripts: Employ long-read RNA sequencing to capture full-length transcript isoforms, alternative splicing events, and gene fusions that short-read methods cannot reliably detect.
- Prioritize Data Quality Over Quantity: Focus on generating highly accurate biological measurements rather than simply accumulating more data, since sequencing errors and fragmented transcripts can mislead AI models.
- Start with Established Biomarkers: Begin by applying long-read diagnostics to fusion genes and other well-characterized biomarkers where clinical utility already exists, then expand to broader functional profiling.
What Does This Mean for Precision Medicine?
The gap between sophisticated cancer drugs and the diagnostics used to prescribe them has been a persistent problem in precision oncology. Researchers have developed many targeted therapies marketed as precise and personalized, but the tests used to decide who receives them often measure indirect signals rather than the actual biology driving treatment response.
Long-read RNA sequencing addresses this gap by providing a more direct window into cancer biology. By reading complete RNA molecules, clinicians can determine not just whether a tumor expresses a target protein, but whether it expresses the specific isoform that a drug is designed to bind. This level of precision allows AI models to make more confident predictions about which patients will respond to which treatments.
Wobble Genomics is initially focusing on fusion gene diagnostics, where RNA transcripts can vary considerably between patients because of different breakpoints and splice patterns. Current panel-based tests may fail to detect clinically relevant fusion variants, but long-read sequencing can capture this diversity. The company's approach demonstrates that this technology can achieve both high sensitivity (detecting rare tumor signals) and high specificity (avoiding false positives), critical requirements for liquid biopsy tests that guide treatment decisions.
As AI becomes more central to drug discovery and patient stratification, the quality of biological data will determine whether these models deliver on their promise. Long-read sequencing provides the rich, multidimensional information that AI needs to uncover meaningful patterns and make predictions researchers can trust.