Logo
FrontierNews.ai

Why AI Companies Like Anthropic Are Betting Big on a Legal Gray Zone

Training AI models on copyrighted books sits in a legal gray zone that favors AI companies more than authors realize. Last year, a federal judge ordered Anthropic to pay $1.5 billion to writers whose works were used to train Claude, but the ruling contained a surprising twist: the judge actually determined that Anthropic's AI training practices were lawful. The penalty came not from the training itself, but from how Anthropic obtained the books, pirating them from illegal online shadow libraries.

What Did the Judge Actually Rule About AI Training?

Judge William Alsup's decision fundamentally reframed how courts should think about AI and copyright. Rather than viewing AI training as copying, Alsup compared it to how human writers learn by reading literature. "Like any reader aspiring to be a writer, Anthropic's LLMs trained upon works not to race ahead and replicate or supplant them, but to turn a hard corner and create something different," the judge wrote. This analogy matters enormously because copyright law hinges on the act of copying, not on simply reading or experiencing a work.

The distinction is crucial for companies building AI models. Copyright law hasn't been updated since 1976, leaving judges to interpret decades-old guidelines when confronting questions that could shape the entire AI industry. Cathy Gellis, an attorney specializing in intellectual property and technology, explained the broader implications: "Copyright law hinges on copying, but it doesn't hinge on using the work or experiencing the work, consuming the work, reading the work".

For Anthropic specifically, the $1.5 billion fine represents a manageable cost relative to its projected revenue. The company is expected to generate roughly $200 billion in annual revenue by 2028, making the settlement a fraction of its anticipated earnings.

How Do Courts Decide What's Legal in AI Copyright Cases?

The legal framework courts use centers on "fair use," a carve-out in copyright law that permits use of copyrighted materials without explicit permission under certain conditions. Fair use protects activities like criticism, parody, education, and commentary on copyrighted works. When deciding whether AI training qualifies as fair use, judges consider several factors:

  • Purpose and Nature: Whether the use serves a transformative purpose or directly competes with the original work's market
  • Amount Used: How much of the copyrighted material was incorporated into the AI training dataset
  • Market Impact: Whether the AI model harms the original creator's ability to profit from their work

Jason Henderson, senior attorney and founder of the IP & Media Practice at JWL International, noted that courts are increasingly focused on competitive intent: "If what you're doing is you're training on somebody's property because your purpose is to directly compete, then the courts will frown on it. If what you're doing is not going to compete, then the courts are tending to find ways that it will be okay".

A case involving Thomson Reuters and Ross Intelligence illustrates this principle. Ross Intelligence built an AI-powered legal research platform by training on Thomson Reuters' content. Judge Stephanos Bibas ruled this was not fair use because Ross's use lacked "a further purpose or different character" than Thomson Reuters's original work, and it directly competed with the company's existing business.

Why Authors Haven't Won in Court Yet

Despite concerns that AI chatbots could compete with human authors by generating synthetic books, this argument has not yet prevailed in any court ruling. Authors would need to demonstrate that AI models trained on their works directly harm their market and serve primarily competitive purposes rather than transformative ones. The current legal landscape suggests this is a difficult threshold to meet.

The uncertainty stems partly from how fragmented copyright law has become in the AI era. Henderson emphasized this challenge: "Everybody is very worried right now because the law is all over the place, and it's because of this question. They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question".

Henderson

What Happens to AI-Generated Content Under Copyright Law?

The copyright questions don't end with training. A separate legal question concerns whether AI-generated content itself can be copyrighted. In the case Thaler v. Perlmutter, a court ruled that works generated entirely by AI cannot be copyrighted, opening new complications. If a work is 100 percent AI-generated, it receives no copyright protection, but determining the exact percentage of AI involvement in a work remains technically and legally murky.

Gellis drew a helpful parallel to existing technology: "If you write your novel in Microsoft Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel. AI is forcing us to look at a whole bunch of decisions that we kind of ignored for a while".

How to Navigate the Current Legal Uncertainty

For AI companies, researchers, and developers working with large language models, the current legal environment requires careful attention to several principles:

  • Source Legitimacy: Obtain training data through legal channels rather than pirated content, as Anthropic discovered when penalized for using shadow libraries despite lawful training practices
  • Transformative Purpose: Ensure your AI application creates something meaningfully different from the original works, not a direct substitute or competitor
  • Market Impact Assessment: Evaluate whether your model could harm creators' ability to profit from their original work before deployment

Most AI companies remain entangled in pending litigation over these issues, meaning definitive legal clarity remains years away. Gellis cautioned that early court decisions are shaping the entire industry even as they remain subject to reversal: "What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it'll take later states of litigation to figure out which one will prevail. But in the meantime, all these decisions are shaping everything that's happening. It would be kind of foolish for the AI companies to ignore them".

Gellis

The Anthropic ruling ultimately signals that AI companies have more legal room to train on copyrighted works than many assumed, provided they obtain the data legitimately and their models serve transformative purposes. However, the landscape remains unsettled, and future court decisions could shift the balance significantly.