Logo
FrontierNews.ai

Fair Use Protects AI Training, But Data Sourcing Creates New Liability: What the $1.5B Anthropic Settlement Reveals

Training an AI model on copyrighted material can qualify as fair use under U.S. law, but acquiring, storing, or mishandling that data creates independent legal violations that fair use does not protect. This distinction, clarified in recent high-stakes litigation, is reshaping how companies approach AI development and vendor selection.

What Changed in the Anthropic Settlement?

In 2025, a federal court in Northern California ruled in Bartz v. Anthropic that training an AI model on lawfully acquired copyrighted books constitutes transformative fair use, a significant legal win for AI developers. However, the same ruling drew a hard line around data handling practices. Anthropic's practice of retaining pirated digital copies of books was not excused by the fair use defense, and the company ultimately settled the case for $1.5 billion, with individual payouts estimated near $3,000 per infringed work.

This outcome signals a critical shift in AI litigation strategy. Rather than fighting over whether training itself infringes copyright, plaintiffs and courts are now focusing on how companies source, acquire, and retain training data. The distinction matters enormously for businesses building or deploying AI systems.

How Does This Apply to the Google Lawsuit?

On July 14, 2026, a coalition of major publishers and authors, including Hachette, Cengage, Elsevier, and novelist Scott Turow, filed a class action lawsuit against Google alleging the company used their copyrighted works to train Gemini models without permission. The complaint goes further than a standard infringement claim: it alleges Google intentionally removed or altered copyright management information (CMI) embedded in the works to obscure the fact that Gemini had been trained on the material.

Copyright management information is the metadata attached to a copyrighted work, such as the author, title, and license terms, that identifies ownership and use restrictions. Stripping or altering this information is legally distinct from ordinary copyright infringement and represents a separate violation that fair use does not excuse.

Steps to Reduce Legal Exposure From AI Training Data

  • Document Data Provenance: Maintain detailed records of where training data comes from, including licensing terms and representations about how datasets were assembled. This documentation becomes critical evidence if your company is later sued.
  • Avoid Retaining Improperly Sourced Material: Do not keep copies of copyrighted works obtained outside of a clear license, even temporarily. Retention itself can create independent liability separate from the training process.
  • Negotiate Vendor Protections: When licensing AI tools from outside vendors, ask pointed questions about their training data sourcing practices and seek contractual indemnification in case the vendor is later found to have used improperly obtained data.
  • Monitor Fair Use Developments: AI copyright law is rapidly evolving. Companies should track ongoing litigation, particularly the UMG Recordings v. Suno case involving music training data, which is scheduled for a summary judgment ruling on fair use in January 2027.

Why Does the Music Litigation Matter?

A parallel case, UMG Recordings, Inc. v. Suno, involves music recordings and is worth watching closely. Discovery in that case surfaced audio fingerprinting evidence showing millions of copyrighted recordings embedded in Suno's training data. A summary judgment ruling on whether AI music training without a license constitutes fair use was originally expected in summer 2026 but has been pushed back repeatedly and is now scheduled for January 2027 before Chief Judge F. Dennis Saylor IV in the District of Massachusetts.

The outcome will likely extend beyond the music industry. Courts often apply fair use rulings across media types, meaning a decision in the Suno case could influence how fair use applies to text, image, and code models as well. This makes the January 2027 ruling a pivotal moment for the entire AI industry.

What Is the Emerging Legal Framework?

These cases are converging on a workable, if still evolving, legal framework: the act of training a model on copyrighted works may be defensible as transformative fair use, but a company's data acquisition and retention practices are judged independently and can create liability even when the training itself would otherwise be protected. In practice, this means the sourcing pipeline of materials used to train large language models (LLMs) is now a distinct area of legal risk that requires its own diligence.

For companies developing proprietary AI models, fine-tuning foundation models on internal or third-party datasets, or licensing AI tools from vendors, the message is clear: fair use may protect your training process, but it will not protect you if your data sourcing practices are questionable. The liability exposure has shifted from the model itself to the supply chain that feeds it.