Logo
FrontierNews.ai

The Hidden Problem Slowing Down AI Drug Discovery: Why Labs Can't Share Their Failures

AI is accelerating drug discovery by predicting which molecules will work before labs test them, but a fundamental data problem is undermining the entire approach: scientists almost never share their failures. Most publicly available datasets focus exclusively on successful results, leaving AI models without the comprehensive understanding needed to make reliable predictions. This bias is forcing researchers to rethink how data flows between computational labs and physical laboratories.

Why Are Failed Experiments So Valuable to AI Models?

Drug discovery has always been expensive and risky. Bringing a new drug to market takes an average of 10 to 15 years and costs anywhere from $1 billion to $2.5 billion, with failure rates exceeding 90%. AI promised to change that by helping researchers identify promising drug candidates faster, eliminating low-quality options before expensive lab work begins.

But here's the problem: AI models trained on incomplete data can't learn what doesn't work. "Most publicly available datasets and scientific publications focus exclusively on positive results," explained Paul Belcher, director of protein research strategy at Cytiva. "No one wants to share their failures. This bias is almost like having one hand tied behind your back. AI models can identify patterns associated with success, but they lack the comprehensive understanding of failures that would make predictions more reliable".

The compounds that don't bind to disease targets, the experiments that produced no useful data, the failed optimization attempts,this information remains buried in lab notebooks and proprietary databases. Without access to these negative results, AI models reach a plateau where they all draw similar conclusions from the same limited public data, producing diminishing returns over time.

How Is This Affecting Drug Discovery Right Now?

The impact is already visible in labs. As AI has accelerated the discovery of potential drug candidates, it has created a new bottleneck: validation. Traditional screening workflows were designed to identify promising molecules at scale using simple yes-or-no tests. But AI-generated candidates are more diverse and complex, requiring detailed characterization and testing.

"AI can increase the number of hits you get and potentially give you better quality hits as well," Belcher noted. "That increases demand for higher-throughput, information-rich technologies to then validate and characterize those hits". Lab teams now face pressure to test, characterize, and purify a growing volume of AI-designed compounds, but they're doing so with AI models that lack critical failure data.

The problem extends beyond incomplete datasets. Data integrity has become a serious concern as generative AI makes fabrication easier. Research by Dutch microbiologist Elisabeth Bik found that almost 4% of biomedical papers contained duplicated or manipulated images, and that was in 2016, before modern AI tools made creating fake data trivial.

What Steps Are Researchers Taking to Close the Data Loop?

  • Building Autonomous Labs: The future of drug discovery involves fully autonomous laboratories that run with minimal human intervention, cycling through prediction, testing, and optimization while feeding results back into AI models to guide the next round of experiments. This closed loop between computational and physical labs could improve success rates of drug candidates entering clinical trials.
  • Implementing Data Integrity Tools: Some vendors are developing solutions to detect manipulated scientific images using secure hash algorithms, the same technology used in blockchain. Publishing houses are beginning to adopt these tools as standard practice to ensure published data is genuine.
  • Creating FAIR Data Infrastructure: Labs are moving toward integrated systems that generate FAIR data (findable, accessible, interoperable, and reusable) at scale. This requires interoperable instruments, highly structured datasets, and information flowing easily between systems. Most labs today still operate with standalone instruments in closed ecosystems where data can't easily be extracted or shared.

The challenge is that most laboratory instruments today are standalone systems. "You can have the best technology in the world, but if it's a closed ecosystem, if the user can't get the data out, it doesn't do any good," Belcher explained.

What Role Does Data Quality Play Beyond Just Quantity?

The shift toward AI-driven drug discovery has exposed a deeper truth: more data isn't enough if the data lacks structure, proper labeling, and diversity. Early AI models were trained on publicly available datasets that weren't built with machine learning in mind, meaning they lack the organization and annotation needed to keep models accurate and free of bias.

Dr. James McDonagh, Principal Applied AI Scientist at ApconiX, emphasized that data curation remains the biggest blind spot in AI drug discovery. "The biggest blind spot behind all of these advances remains careful data curation, annotation, and collection," he stated. "It is not just about data scale, but quality that is critical here. Accessible high quality datasets are the fuel of AI methods. Architecture matters, but high-quality, accessible and abundant data is also absolutely critical".

This insight applies across different AI approaches. While graph neural networks and foundation models have transformed fields like protein structure prediction, classical machine learning methods still outperform newer approaches in certain scenarios. "Classical machine learning models are still very useful, especially for small datasets, which are common in early discovery," McDonagh noted. "Classical models often have fewer parameters and can generalize better in small chemical spaces than larger deep learning models, which may overfit in such spaces".

What's Changing in How Companies Approach AI Drug Discovery?

Major pharmaceutical companies are recognizing that AI-enabled discovery requires more than just better algorithms. Harbour BioMed, a global biopharmaceutical company, recently announced a strategic collaboration with Sinopharm to establish an innovation consortium focused on advancing antibody therapeutics. The partnership leverages Harbour BioMed's fully human antibody technology platform and AI-enabled drug discovery capabilities alongside Sinopharm's clinical development, manufacturing, and commercialization expertise.

"Harbour BioMed is committed to becoming a leading platform-driven biopharmaceutical group," said Dr. Jingsong Wang, founder, chairman and CEO of Harbour BioMed. "Leveraging our globally differentiated Harbour Mice fully human antibody technology platform and continuously expanding AI-enabled antibody discovery capabilities, we have established strategic collaborations with leading global pharmaceutical companies, including AstraZeneca, Bristol Myers Squibb and Pfizer".

Dr. Jingsong Wang, founder, chairman and CEO of Harbour BioMed

These partnerships reflect a broader shift: AI drug discovery works best when computational capabilities are integrated with manufacturing scale, clinical expertise, and access to diverse patient populations. No single company can optimize the entire pipeline alone.

The path forward requires a fundamental change in how the scientific community treats data. Until researchers begin sharing negative results, publishing failed experiments, and building integrated lab systems that generate high-quality, structured data, AI models will continue operating with one hand tied behind their back. The future of faster, cheaper drug discovery depends not just on better algorithms, but on better data practices.