Logo
FrontierNews.ai

Why AI Can Generate Millions of Scientific Discoveries But Can't Verify Them

AI has solved the generation problem in science, but created a new crisis: verification. Systems like AlphaFold and GNoME now produce candidate structures, hypotheses, and materials faster than any laboratory can test them. The scarce resource in modern science is no longer ideas,it's confirmed results. This asymmetry between cheap generation and expensive verification is reshaping how researchers approach AI-driven discovery.

How Has AI Changed the Speed of Scientific Discovery?

The numbers are genuinely exceptional. AlphaFold, DeepMind's protein structure prediction system, jumped from roughly 40% accuracy at scientific competitions before its release to nearly 90% accuracy by 2020, crossing a threshold widely treated as competitive with experimental methods. The system has since been adopted by more than 3.4 million researchers across 190 countries, with over a million users in low- and middle-income countries. AlphaFold 3, released in 2024, extended the system to predict protein-ligand, protein-DNA, and protein-RNA complexes, posting roughly a 50% accuracy improvement over the best physics-based methods.

GNoME, Google DeepMind's graph neural network for materials discovery, predicted 2.2 million new crystal structures in a single effort. DeepMind described this as roughly 800 years of accumulated materials knowledge, delivered as a database. The system identified 528 candidate lithium-ion conductors, roughly 25 times more than a previous landmark study, and 52,000 layered, graphene-like compounds.

But here is where the story gets complicated. Of GNoME's 380,000 most-stable crystal predictions, only 736 have been independently synthesized by laboratories worldwide. That is about 0.2%. GNoME did not discover ten times more materials that exist; it discovered ten times more hypotheses about materials that might exist, the overwhelming majority of which have never been made outside a simulation.

Where Is the Verification Bottleneck Actually Happening?

The AI discovery loop operates in three distinct layers, and the bottleneck sits between them. Layer 1 includes prediction engines like AlphaFold and GNoME that turn specifications into candidates. Layer 2 involves hypothesis agents, including systems like RLVR (reinforcement learning verification and reward modeling), that sequence reasoning about those candidates by reading literature, proposing hypotheses, and planning experiments. Layer 3 executes experiments with liquid handlers, solid-state synthesizers, and cloud labs, feeding measurements back into the models.

Throughput above the verification gate moves at the speed of computation: days. Throughput below it crawls: months to years, sometimes longer if a "verified" result turns out to need re-verification. A widely cited autonomous-lab result that claimed 41 new materials has been challenged by independent chemists who argue the synthesis was not verified correctly at all.

Even in drug discovery, where AI has shown genuine promise, the pattern holds. AI-discovered drug candidates clear Phase I trials at roughly 80 to 90%, versus a 40 to 65% historic norm. But Phase II success sits around 40%, statistically close to the conventional baseline. AI is improving molecule engineering, not target biology. As of September 2026, Insilico Medicine's rentosertib became the first drug with both an AI-identified target and an AI-designed molecule to enter Phase III trials, a genuine milestone, though still short of approval.

What Are the Practical Limitations of Current AI-Generated Predictions?

AlphaFold's success masks several important caveats for researchers building on these structures. The system shows a measurable rate of chirality violations and occasional overlapping atoms in multi-chain assemblies, the kind of geometric error a bench chemist would catch immediately but a computational pipeline might not. The model's accuracy also drops on test sets published after its training cutoff, a pattern more consistent with pattern memorization than fully generalizable modeling.

Additionally, AlphaFold's confidence score estimates coordinate accuracy, not functional or binding correctness, and shows weak correlation with experimentally measured binding affinities. The model also hands researchers a single conformation of a system that, in reality, usually moves and changes shape. One published TAAR1-agonist screen using AlphaFold-derived models reported roughly a 60% hit rate versus 22% for classical homology modeling, but ligand-bound conformations still need experimental confirmation before researchers make structure-activity decisions on their basis.

How Should R&D Leaders Integrate AI Into Their Workflows?

  • Screening and Triage: Buy AI where a wrong answer is cheap to catch. Use AI systems for initial screening and structure triage, where errors can be quickly identified and corrected without major cost or time investment.
  • Synthesis Planning: Deploy AI for synthesis planning and molecular design, where computational predictions can guide laboratory work but do not replace it.
  • Human Validation Loop: Keep humans and wet-lab validation in the loop where a wrong answer costs a decade. Every flagship autonomous-lab demonstration so far runs on curated tasks with a human checking the output, and none has independently delivered a validated, field-accepted discovery from start to finish.

The practical rule is straightforward: AI excels at generating hypotheses and narrowing the search space. It struggles with the expensive, time-consuming work of confirming that those hypotheses reflect reality. Autonomous labs and multi-agent research systems compress design-make-test-learn loops, but they still require human oversight at critical verification points.

The contest over whether GNoME's predictions are real illustrates the broader thesis perfectly. In November 2023, two papers landed in Nature on the same day. The first described GNoME's 2.2 million predictions. The second described A-Lab, a Lawrence Berkeley National Laboratory facility where AI-guided robots tried to make 58 of those predicted materials and reported success with 41 of them over 17 days of continuous operation. Sit with that ratio: 800 years of candidate knowledge, produced almost overnight, verified at a rate of about 2.4 materials per day. That asymmetry, nearly-free generation against expensive verification, is the single most important fact about AI in science in 2026.