The AI Washing Problem: How Digital Health Buyers Are Getting Fooled by Fake AI Claims
When health tech companies claim to have proprietary artificial intelligence, they may actually be using basic analytics wrapped around third-party tools, inflating valuations in the process. A new legal analysis from Jones Day reveals how "AI washing" has become a widespread problem in digital health mergers and acquisitions, where target companies disguise basic analytics as proprietary medical diagnostics to command premium prices.
What Exactly Is AI Washing in Digital Health?
AI washing occurs when digital health companies exaggerate their actual artificial intelligence capabilities to inflate valuations during acquisition discussions. The challenge for buyers is that these assets present unique obstacles for standard due diligence: opaque model architectures, unverifiable training datasets, and an unsettled intellectual property landscape where copyright exposure, open-source license considerations, and data provenance issues remain largely unquantified.
The underlying technology cannot be audited line by line like regular software because its capabilities emerge from numerical parameters derived through training on data of varying provenance and legality. This means old review checklists cannot account for these risks, and regulators are rapidly adding compliance dimensions that increase transactional complexity.
How Can Buyers Identify Real AI Versus Inflated Claims?
To spot AI washing, buyers must first address a threshold question: what exactly is being purchased? Defining the target company's competitive advantage, often called its "moat," is essential. The biggest valuation question is whether the technology represents a proprietary foundation model or merely a basic "wrapper" on a third-party application programming interface, or API.
A digital health company commanding a premium "AI company" price may only possess engineered prompts and an API connection to a commercial foundation model, or a thin integration layer built on third-party machine learning services. This setup offers minimal defensibility and significant platform dependency risks. Conversely, a proprietary foundation model trained from scratch on demonstrably clean data, with genuinely novel architecture, represents a highly defensible asset with significant enterprise value.
Steps to Conduct Effective AI Due Diligence
- Evaluate Replicability: Determine whether a well-funded competitor could rebuild the core functionality within months using publicly available foundation models, which would indicate the technology lacks true defensibility.
- Trace Data Provenance: Meticulously trace the provenance of every material training dataset, since mere access to health insurance data, wearable sensor streams, or electronic health record APIs through health system partnerships does not grant the commercial right to train models on that data.
- Assess Proprietary Data Assets: Identify whether the true value lies in proprietary data assets such as patient interaction logs, labeled clinical datasets, and curated health knowledge bases, provided that clear chain of title and compliant collection practices can be demonstrated.
- Quantify Risk Rather Than Eliminate It: Embrace a paradigm shift from issue elimination to risk quantification, developing a sufficiently informed assessment of the risk magnitude, probability of materialization, and potential liability exposure.
Why Data Provenance Matters More Than You Think
The most challenging phase of AI due diligence involves identifying risks that could create material post-closing liability. In health care AI, privacy, data security, and regulatory compliance are especially critical given sector-specific regimes such as the Health Insurance Portability and Accountability Act, or HIPAA.
However, conducting a complete audit of AI training data is not feasible given normal transaction timelines and resource constraints. These models may have trained on datasets from hundreds or thousands of sources, including copyrighted content potentially ingested without authorization, personal data processed without valid legal basis under applicable privacy regimes, and contractually restricted data subject to use limitations or confidentiality obligations.
Using unauthorized or improperly de-identified clinical data creates what legal experts call a "fruit of the poisonous tree" scenario: even if the resulting algorithm performs well, the tainted data provenance can give rise to massive infringement claims, regulatory disgorgement actions, and in the most severe cases, the forced destruction of the algorithm itself.
Can AI Assets Be Salvaged After Problems Are Discovered?
Discovering material issues during due diligence does not have to kill a digital health transaction. Unlike traditional software, where affected code may require a total rewrite, machine learning systems can often be remediated in a more technically and economically feasible manner because models are designed for iterative improvement.
For example, if a medical AI is trained on legally problematic data, it can be retrained using clean datasets; algorithmic bias can be mitigated through targeted fine-tuning; and open-source issues in certain model components may be isolable and replaceable. Acquirers have several practical options to save the transaction and protect their investment.
Ring-fencing can exclude problematic assets from the initial transfer, with options for post-closing acquisition once retraining is completed and validated. Additionally, post-closing retraining covenants can obligate the seller to undertake defined remediation within specified periods, with performance benchmarks, testing procedures, and milestone-based escrow releases providing accountability.
After closing, buyers must be prepared to integrate the acquired AI assets into their existing compliance and governance infrastructure. This requires coordination between outside counsel, in-house counsel, and technical, compliance, and business stakeholders from the earliest diligence stages. Buyers must plan for regulatory requirements, including ongoing monitoring, testing, and validation procedures consistent with applicable AI governance frameworks, and identify any required regulatory filings or approvals to deploy AI in regulated health care business lines.
Ultimately, perfection is unrealistic in AI diligence, but informed, structured pragmatism can produce outcomes that address each party's interests. By deconstructing the AI asset, deploying tailored contractual protections, treating remediation as integral to deal design, and planning for post-closing regulatory compliance, buyers can navigate the uncertainty and acquire truly valuable digital health assets without sacrificing deal momentum or commercial viability.