Logo
FrontierNews.ai

AI Drug Discovery Just Hit a 100x Breakthrough. Here's Why Physics Matters More Than Pure Machine Learning

A San Francisco startup has achieved a major breakthrough in AI-powered drug discovery by combining physics-based simulations with machine learning, delivering hit rates nearly 100 times higher than conventional AI-only approaches. Deep Origin announced that its AI drug discovery framework achieved a 31% hit rate for a cancer immunotherapy enzyme, a dramatic improvement over the roughly 1% success rate typical of pure machine learning virtual screening methods.

Why Do AI-Only Drug Discovery Models Fail on New Targets?

The challenge in computational drug discovery has long been a trade-off between accuracy and generalization. Modern deep learning models like AlphaFold and Boltz-1 are trained on vast repositories of existing molecules, which can boost their predictive accuracy to over 80% on familiar compounds. However, when these models encounter novel compounds they've never seen during training, their accuracy deteriorates significantly.

This limitation has plagued the field for decades. Virtual screening has been used for more than 40 years, yet it has rarely delivered an actual drug to patients. The industry learned to accept a low accuracy ceiling, treating false positives as an inevitable cost of the process.

"Virtual screening has been widely used for more than 40 years, yet it has rarely delivered an actual drug. The field learned to live with a low accuracy ceiling. We set out to change that by pairing physics with machine learning, ensuring our method holds its accuracy on novel targets where training data alone runs out," stated Garegin Papoian, chief scientific officer of Deep Origin.

Garegin Papoian, Chief Scientific Officer at Deep Origin

How Does Deep Origin's Physics-Plus-AI Approach Work?

Deep Origin's breakthrough relies on two complementary products that work together to overcome the generalization problem. The company's framework combines learned components with interpretable physics models, placing each where it adds the most value.

  • DODock: An AI-powered and physics-based molecular docking engine that uses a diffusion model paired with physics-based search and refinement built on a Vina-style energy function, allowing the system to propose and rank molecular poses before a compact physics model refines them based on how molecules physically interact.
  • DOScore: An engine for estimating thermodynamic binding affinities across ultra-large chemical spaces, helping predict how strongly molecules bind to target proteins.
  • Physics-First Refinement: Once machine learning proposes candidate poses, a physics-based model refines them using first principles rather than relying on resemblance to training data, enabling the system to maintain accuracy on novel targets where AI-only approaches lose precision.

The key insight is that physics allows the system to build on fundamental principles rather than on prior data. This distinction is critical because it enables the software to maintain its accuracy on novel targets, where AI-only approaches struggle. Deep Origin says that rebuilding established C++ and Fortran-based physics engines to run efficiently on GPUs represents a technical hurdle that most startups lack the expertise to overcome.

What Real-World Results Did Deep Origin Achieve?

Deep Origin tested its framework against four challenging biological targets: CD73 (nucleotidase), IRAK4 (kinase), Factor XIa (protease), and IL-17A (a protein-protein interaction target). The company is publishing detailed wet-lab validation results in BioRxiv, demonstrating that its approach discovers entirely new chemical matter rather than minor variations on existing drugs.

The most striking result came from the CD73 target, which is important in slowing cancer's growth. Deep Origin achieved a 30.6% hit rate in finding ligands that pair with the enzyme, approximately 100 times better than pure AI and machine learning-based approaches. Across all four targets, the discovered compounds demonstrated high scaffold novelty, with a Tanimoto similarity of 0.22 to 0.27 relative to known binders, indicating genuinely novel chemical structures.

"What matters is how a screen behaves on a target no one has solved yet, so that is where we tested our model with strict splits and cases built to make it fail. How our model performs on those rigorous tests, not on inflated benchmarks, is what dictates how many molecules brought to the bench are likely to be real," stated Michael Antonov, CEO of Deep Origin.

Michael Antonov, CEO at Deep Origin

How Does Deep Origin Prevent Overfitting and Data Leakage?

A critical aspect of Deep Origin's validation methodology is its strict approach to preventing data leakage, a common problem in machine learning where models inadvertently memorize test data rather than learning generalizable patterns. The company uses a "strict double-similarity filter" that removes any protein sharing 30% or more sequence identity and any ligand exceeding 0.4 Tanimoto similarity.

This rigorous approach allows Deep Origin to successfully run virtual screens against novel biological targets without existing crystallographic data or known small-molecule binders. The company claims it maintains 50% to 89% pose accuracy on these challenging cases, where pure AI and machine learning approaches failed entirely.

Deep Origin was founded in 2022 by Michael Antonov and Garegin Papoian and has raised $82 million in funding. The company's long-term vision is to fundamentally change how the pharmaceutical industry discovers new drugs by proving that physics-informed machine learning can deliver real, testable compounds rather than theoretical predictions.