How AI Drug Discovery Benchmarks Are Being Rigged,and What One Company Is Doing About It
Most AI drug discovery benchmarks are contaminated with test questions that models have already seen during training, making their high scores meaningless in real-world drug development. Insilico Medicine, a Hong Kong-listed biotech company, is now opening its internal evaluation framework to outside AI developers to address this problem and reshape how pharmaceutical companies assess AI vendors.
Why Are Public AI Benchmarks Failing Drug Discovery?
The problem sounds deceptively simple but has major consequences. When AI models are trained on massive datasets, they sometimes encounter variations of the same test questions used in public benchmarks. This means a model can score 90% on a chemistry benchmark not because it understands drug synthesis, but because it has essentially memorized similar problems during training. For pharmaceutical companies deciding whether to invest in AI-powered drug discovery, this distinction is critical.
Insilico argues that its own datasets come from proprietary validated programs with explicit decontamination steps designed to prevent memorization. The company has used generative AI across its pipeline since its founding, nominating 31 preclinical candidates in six years and compressing the typical 2.5 to 4 year timeline to preclinical nomination down to roughly 12 to 18 months. Its lead program, Rentosertib, is currently in Phase III clinical trials for idiopathic pulmonary fibrosis.
What Does the New Benchmark Actually Test?
The Drug Discovery and Development Benchmark as a Service, or DDD BaaS, launched on July 30 and covers two evaluation suites designed to measure whether AI models can actually navigate real drug discovery challenges.
- Drug Discovery Foundations: Comprises more than 300 tasks spanning disease biology, molecular property prediction, retrosynthesis, structure-based drug design, and clinical development stages.
- Drug Candidate Essentials: Tests a model's ability to run an end-to-end discovery program, from hit identification through to preclinical candidate nomination, simulating the full workflow.
- Decontamination Protocol: Uses proprietary datasets with explicit steps to prevent models from scoring well simply because they have seen similar questions before.
The distinction matters because a model that performs consistently against Insilico's closed dataset represents a fundamentally different proposition than one that scores well on public benchmarks but struggles with novel synthesis planning against proprietary targets.
How to Evaluate AI Drug Discovery Tools for Your Organization
- Request Proprietary Benchmark Results: Ask vendors whether their AI has been tested on decontaminated datasets that prevent memorization, not just public benchmarks where high scores may be misleading.
- Assess End-to-End Capability: Evaluate whether the AI can handle the full drug discovery pipeline from target identification through preclinical candidate nomination, not just isolated tasks like molecular property prediction.
- Verify Real-World Timelines: Look for evidence that the AI has actually compressed drug discovery timelines in practice, such as preclinical candidates nominated in 12 to 18 months rather than the traditional 2.5 to 4 years.
- Test Against Your Data: Use evaluation frameworks like DDD BaaS that allow you to benchmark models against your own proprietary targets and workflows rather than relying solely on vendor claims.
The broader context is a competitive market where numerous AI drug discovery companies are vying for pharmaceutical partnership deals, and differentiation on benchmark performance has become as much a marketing exercise as a scientific one. An evaluation framework that makes those scores harder to game could reshape how pharma companies assess AI vendors or, if Insilico's dataset construction holds up to scrutiny, give the company a second business line alongside its own drug pipeline.
The DDD BaaS service is now available at dddbench.insilico.com to any organization developing frontier AI for drug discovery or using foundation models in research workflows. Pricing and tiering details were not disclosed in the announcement, but the platform represents a significant shift toward transparency in how AI performance is measured in one of biotech's most competitive spaces.
For pharmaceutical companies and AI developers, the message is clear: benchmark scores alone tell an incomplete story. The real test is whether an AI system can solve novel problems it has never encountered before, not whether it can recognize variations of questions it has already seen.