Logo
FrontierNews.ai

The Real Reason AI Labs Are Losing Copyright Cases: It's Not About Training

The lawsuits reshaping AI training data practices aren't really about whether AI can learn from copyrighted works,they're about whether companies obtained that data legally in the first place. As courts rule on major copyright cases throughout 2026, a pattern is emerging that most coverage misses: fair use protections for training don't shield labs from liability over data provenance, and that distinction is costing companies billions.

Why Are AI Labs Losing If Training Is Fair Use?

Two federal judges in California ruled in mid-2025 that training large language models (LLMs), which are AI systems trained on vast amounts of text to predict and generate language, on copyrighted books can qualify as fair use. In Bartz v. Anthropic, Judge William Alsup called training "transformative, spectacularly so." Meta won a similar fair use argument in Kadrey v. Meta. Yet these victories came with massive costs and ongoing exposure.

Anthropic's $1.5 billion settlement in Bartz v. Anthropic illustrates the disconnect. The court ruled that training itself was fair use, but Anthropic still paid roughly $3,000 per work across about 500,000 books because it had downloaded much of its library from pirate sites like Library Genesis. The court treated acquisition as a completely separate legal question from training. Fair use covered how the company used the data. Nothing covered how it obtained the data. Final approval came in July 2026, and the settlement also required Anthropic to destroy the pirated datasets.

Meta faced a similar split outcome. The company won on fair use against authors who sued, but claims over its torrenting activity, specifically "seeding" pirated files back to other users, remain active in 2026. Again, the exposure that survived is about how the data was obtained and handled, not the training itself.

How Is the Legal Landscape Splitting Across Different Courts?

The current state of copyright law for AI training is fragmented. A third court went the opposite direction from the California judges. In Thomson Reuters v. Ross Intelligence, the court found that copying Westlaw headnotes to train a competing legal research tool was not fair use, largely because the output competed directly with the source material. That case is now on appeal at the Third Circuit.

Getty Images sued Stability AI in the UK, but that case largely failed in late 2025 on jurisdictional grounds because the training happened outside the UK. Getty's US case continues. Disney and Universal are seeking up to $150,000 per work against Midjourney for willful infringement over character images, though Disney simultaneously signed a licensing deal with OpenAI that lets Sora generate video with Disney characters. The music industry initially sued Suno and Udio for training on recordings without permission, then settled in late 2025 by converting the disputes into licensing agreements.

The newest major case arrived in July 2026. Publishers sued Google, alleging the company trained Gemini on books that publishers had provided to Google Books for search snippets only. The complaint cites an internal Google document warning that training on copyrighted books could mean "$10Bs-$100Bs in potential fines." Discovery will test that claim.

What's Changing the Legal Ground Under AI Training Data?

Three patterns are reshaping how courts and companies approach AI training data. First, courts consistently separate two distinct legal questions: whether training on a work is fair use, and whether the copy you trained on was lawfully acquired. A lab can win the first and still lose the second. That is exactly what happened to Anthropic, and it is the theory keeping Meta's torrenting claims alive.

Second, fair use analysis weighs whether copying damages the market for the original work, including the market for licensing it. In 2023 there was barely a licensing market for training data, which made market harm hard to prove. In 2026 there is one. Disney licenses to OpenAI. Warner and Universal license to Suno and Udio. Every major lab has signed data licensing deals. This creates an uncomfortable feedback loop for anyone relying on fair use: each new licensing deal makes the licensing market more real, and the more real that market is, the weaker the fair use defense becomes for training on unlicensed data. The legal ground under scraped data is not holding steady. It is eroding.

Third, even the labs that won spent years in discovery, produced internal documents and output logs by the millions, and absorbed legal costs and headlines. For a frontier lab, the cost of defending a data provenance case now rivals the cost of just licensing the data. It is one of the forces reshaping where AI labs source training data.

How to Evaluate Training Data Sources for Legal Risk

  • Document the Chain of Custody: Where did each asset come from? Not the category, the chain. Who created it, who holds the rights, and what did they agree to? "Publicly available" is not a provenance answer, and the Google case shows that even data acquired legitimately for one purpose can create exposure when used for another.
  • Verify Licenses Cover AI Training: Older content licenses often cover distribution or display but say nothing about model training. A rights-cleared dataset means the rights holder explicitly agreed to training use, not just general use.
  • Require Documentation Per Asset: If a dispute or a due-diligence review ever happens, you want a paper trail per asset, not assurances. Providers should be able to document the provenance of each piece of training data.
  • Assess Data Exclusivity: Is the data exclusive or resold everywhere? Beyond the legal question there is a model quality question. Data that every lab already has moves no benchmarks and offers no competitive advantage.

The lawsuits are not saying copyrighted material is off limits for training. They are saying unaccounted-for material is a liability. The distinction matters because it shifts the burden from defending fair use in court to documenting legitimate acquisition upfront.

Twelve consolidated cases, including The New York Times suit, are moving through discovery in the Southern District of New York against OpenAI. Courts have ordered OpenAI to produce tens of millions of output logs. Whatever the outcome, it has already shown that a lab's data practices will be examined in granular detail once litigation starts.

The 2026 case tracker reveals that the real battle over AI copyright is not about whether training is transformative. Courts have mostly agreed it is. The battle is about where the data came from, who had the right to share it, and whether that right explicitly included AI training use. Companies that can answer those questions with documentation will face far less exposure than those relying on fair use arguments alone.