Logo
FrontierNews.ai

Why Smarter AI Models Won't Save Drug Discovery Without Better Medical Language

The bottleneck in AI-powered drug discovery isn't computation or model sophistication; it's the medical language that AI systems use to understand patient data. When pharmaceutical researchers try to identify patients for clinical trials or analyze real-world evidence, they're working with data coded in billing-era terminology systems like ICD-10 that were never designed to capture the clinical nuance modern precision medicine requires. The result is that AI models, no matter how advanced, inherit incomplete or flattened information from the start.

Why Are Standard Medical Codes Failing Modern Drug Research?

Most AI systems used in life sciences analytics are trained on standardized code sets including ICD-10, SNOMED CT, and LOINC. These standards were built to ensure consistency across billing, documentation, and interoperability in healthcare. But they create a fundamental problem for researchers: the clinical meaning that matters for drug discovery gets lost in translation.

Consider a researcher trying to identify patients with mild, moderate, or severe psoriasis for a clinical trial. The standardized code might simply say "L40: psoriasis." But the clinician's actual documentation reads "psoriasis, moderate severity, with comorbid rheumatoid arthritis." That specificity determines trial eligibility, treatment response, and outcomes. The nuance exists at the point of care, only to be flattened the moment it enters the billing layer.

This creates three recurring failure points across pharmaceutical research:

  • Rare Disease Invisibility: Patients with rare diseases are effectively invisible in datasets because there is no precise code to identify them. A 2024 study in Orphanet Journal of Rare Diseases found that just 34% of 454 rare diseases could be specifically coded in ICD-10-GM, meaning two-thirds of rare disease patients cannot be reliably identified for research.
  • Severity Collapse: Clinically meaningful distinctions between early-stage and late-stage disease, or mild versus moderate versus severe presentations, often collapse into a single standardized category, making it harder to isolate the patient subgroups researchers actually need.
  • Hidden Precision Medicine Data: Molecular subtypes, progression markers, and clinically meaningful modifiers often hide in the free-text sections of electronic health records (EHRs). A study in JMIR Medical Informatics showed that only 13% of extracted concepts from patient records showed any overlap between structured codes and free-text notes, meaning the vast majority of clinical information exists in one form or the other, not both.

What Does This Mean for AI Drug Discovery in Practice?

The consequence is predictable and widespread. Models that look sophisticated in demonstrations fail at the questions clinicians and researchers actually need answered. When a researcher wants to identify patients with mild, moderate, or severe disease for a trial, the AI system trained on flattened billing codes cannot reliably distinguish between them. When a team tries to build a cohort of patients with a specific rare disease, the data simply doesn't contain the precision needed.

A study of vaccine administration published in Frontiers in Digital Health illustrates the scale of this problem. Using natural language processing (NLP), a technique that allows AI to understand human language, researchers found that NLP analysis of unstructured data in patient records led to a 16.8% increase in the identification of vaccine administrations compared with using structured data alone. This gap reveals how much clinically important information is being lost when researchers rely only on standardized codes.

The industry response to AI limitations has largely been to chase larger models and larger training sets. But this approach misses the root problem. Every clinically meaningful AI decision depends on something more fundamental: whether the data beneath the model can accurately express what a physician meant when documenting a patient encounter.

How to Build Better Terminology for Healthcare AI Systems

Fixing this problem requires rethinking the terminology layer that AI systems reason over. Rather than relying on billing-derived codes, pharmaceutical leaders should demand terminology infrastructure with three defining characteristics:

  • Clinical Curation: The terminology must be curated and clinically validated by physicians, terminologists, and subject matter experts, not assembled passively or machine-learned by brute force from billing data alone.
  • Provenance and Traceability: Researchers should be able to trace concepts back to trusted clinical sources and understand exactly why a patient was included in a cohort, ensuring transparency and reproducibility.
  • Connected Knowledge Graphs: The terminology must function as a connected, evolving knowledge graph rather than a static lookup table, allowing researchers to reason consistently across EHR data, registries, claims, and literature without losing clinical fidelity.

This is not fundamentally a model problem; it is a precision problem. The infrastructure layer that many healthcare AI deployments are still missing is not about computation or algorithmic sophistication. It is about ensuring that the clinical meaning preserved in patient records survives the journey into the AI system intact.

What Should Pharma Leaders Do Right Now?

For life science research leaders, the implications are immediate and actionable. The conversation around trustworthy healthcare AI is increasingly less about the model itself and more about the clinical intelligence beneath it. Before investing in another model upgrade or larger training dataset, leaders should audit the data layer underneath their existing AI systems.

Ask what terminology and clinical content the model is reasoning over, how much specificity survives end-to-end, and where that specificity gets lost. Demand provenance so every cohort definition and evidence synthesis can be traced back to clinical source material. The marginal value of a more advanced model, given inadequate terminology, is relatively small. The marginal value of better terminology underneath a competent model is enormous.

The real bottleneck in AI-enabled drug discovery is not whether the next generation of models will be smarter. It is whether the language used to describe patients and diseases will finally be precise enough for AI to reason over accurately. Until that changes, even the most sophisticated AI systems will continue to inherit the same flattened, incomplete view of patient data that has always limited pharmaceutical research.