Logo
FrontierNews.ai

The DOJ Just Signaled How AI Companies Will Pay for Training Data. Here's What It Means for Your Business.

The Department of Justice has signaled that Congress, rather than courts, should set rules for how AI companies pay for copyrighted training data. In a filing submitted to federal court this month, the DOJ argued that forcing AI developers to license material could let large media companies set prices only major tech firms could afford, potentially shutting out smaller competitors entirely.

Why Is the DOJ Weighing In on an AI Copyright Case?

The DOJ filed what's called a "statement of interest" in the multidistrict litigation brought by The New York Times, New York Daily News, and the Center for Investigative Reporting against OpenAI and Microsoft in the US District Court for the Southern District of New York. This type of filing carries significant weight because it represents the federal government's institutional position on how the law should develop.

The brief was signed by associate attorney general Stanley E. Woodward Jr., assistant attorney general Brett Shumate, and senior counsel Michael Weisbuch. Their core argument: training an AI model on copyrighted text constitutes "fair use" because the use is transformative, and ruling otherwise would hamper scientific progress and innovation, which copyright law is designed to protect.

But the filing makes a more strategic move. The publishers argue that free AI training destroys a licensing market they could otherwise build. The DOJ flipped that argument: requiring licenses would let a handful of large media companies set prices that only the biggest AI developers could afford, protecting incumbents on both sides and shutting out smaller competitors.

How Does This Compare to the Anthropic Settlement?

The timing of the DOJ's brief is notable. It arrived five weeks after Anthropic finished paying $1.5 billion to settle a similar copyright dispute, the largest publicly reported copyright recovery in history. However, the two cases present vastly different risk profiles.

Anthropic's settlement stemmed from litigation over its downloading of books from pirate sites. The court had already found infringement arising from that piracy, though the settlement did not resolve whether AI training itself constitutes infringement when underlying works were lawfully acquired.

The New York Times case raises both acquisition and downstream-use questions. The plaintiffs allege that OpenAI acquired copies of their articles without authorization, including material obtained from behind paywalls or in violation of applicable terms or industry norms. This distinction matters enormously for how companies should structure their AI training practices and vendor agreements.

What Should Companies Do Right Now?

Lawyers advising on AI vendor contracts, data licensing agreements, and technology mergers and acquisitions shouldn't wait for a final court ruling or congressional action. The Anthropic settlement and OpenAI litigation together sketch two vastly different risk profiles for the same underlying activity: training a model on someone else's content.

  • Provenance Due Diligence: Vendor questionnaires for any company that trains or fine-tunes models should ask how training data was sourced, whether the vendor can identify the origin of material in its training data, and whether any sourcing resembles the bulk downloading at issue in the Anthropic case.
  • Distinguish Risk Types in Contracts: Representations and warranties in AI vendor agreements should be drafted to distinguish provenance risk (how data was obtained) from use risk (how data is used to train a model) rather than folding both into a single generic intellectual property representation.
  • Account for Multiple Outcomes in Indemnification: Indemnification provisions negotiated before litigation concludes should account for the possibility that a favorable outcome for OpenAI narrows fair use risk without eliminating provenance risk, and vice versa if the court rules the other way.

The distinction between these two risk categories is critical. A company can have clean provenance and still face significant fair use risk, or have messy provenance while avoiding per-work damages if it can show something closer to authorized access.

What Does This Mean for Content Licensing Deals?

Publishers and data providers negotiating licensing deals today face a mirror version of the same problem. A publisher or data provider is now negotiating against a backdrop where the DOJ has told a federal court that requiring licenses at all may be bad policy.

That doesn't make licensing deals worthless, but it changes their leverage calculus, particularly for smaller content owners who lack the negotiating power of The New York Times. Deal terms that assumed courts or Congress would eventually mandate licensing may need to be rethought in light of the DOJ's position.

The case is still in discovery and a ruling isn't imminent, but an appeal seems likely given the stakes involved. In the meantime, the DOJ's filing serves as the clearest available signal of how this question may get resolved long before Congress or the courts settle it for good.