Why a Fair Use Ruling Against AI Could Make Training Data Unaffordable for Everyone
A consolidated copyright lawsuit against OpenAI and other AI developers could fundamentally change how generative AI models are built, potentially making them far more expensive and less accessible to schools, startups, and smaller organizations. The central legal question centers on whether AI companies can use copyrighted material under "fair use," a doctrine that has protected innovation in the United States for over 150 years.
The case, brought by dozens of authors including George R.R. Martin and major news organizations, is expected to be argued before a federal judge in New York City early next year. At stake is whether companies like OpenAI can train their large language models (LLMs), which are AI systems designed to understand and generate human language, on copyrighted books, articles, and other published works without permission.
What Is Fair Use, and Why Does It Matter for AI?
Fair use is a legal principle written into the Copyright Act 50 years ago that allows copyrighted works to be used without permission if the new use is "transformative," meaning it creates something fundamentally different from the original. The doctrine has long protected authors, journalists, and creators who build on existing works to create new ones.
AI developers argue that training models on copyrighted material is a textbook example of fair use. When an LLM learns from a book or article, the material isn't copied, stored, or resold. Instead, it's transformed into patterns and knowledge that power a new kind of tool, not a substitute for any specific book or article. The developers also cite decades of Supreme Court precedent suggesting that makers of multi-purpose products aren't responsible when someone else uses that product to violate copyright law.
Plaintiffs, however, are asking courts to force AI companies to pull their existing models off the market and retrain them using only licensed material. This would represent a dramatic narrowing of fair use specifically for AI training.
How Would Restricting Fair Use Reshape the AI Landscape?
If courts dramatically limit fair use for AI training, the consequences would ripple across every sector that relies on generative AI, from classrooms to hospitals to startup incubators. The impact would unfold in several interconnected ways:
- Training Costs Would Skyrocket: Today's leading AI models require enormous amounts of data during training. Without fair use, companies would need to buy or license that material. The volume required is astronomically large, making it prohibitively expensive for startups and smaller developers to compete.
- AI Would Become Less Accessible: Higher development costs would translate directly to higher prices for users. Schools, libraries, nonprofits, and governments of all sizes would face steeper subscription costs and lower usage limits. The largest corporations would still afford premium AI systems, but smaller institutions would have fewer choices.
- Scientific Progress Would Slow: Specialized research models built from general-purpose AI systems would become more difficult and expensive to develop. Researchers would lack affordable tools, not because of a shortage of ideas, but because of cost barriers.
- Knowledge Would Become Fragmented: If every copyright owner could decide whether their works were used in AI training, future models would become patchworks with a fraction of their current power. Some models might include certain newspapers but not others, contemporary books but not historical archives, and major publications but not local reporting.
What Are the Real-World Consequences of a Fragmented Knowledge Base?
Today's best AI models owe their capabilities directly to training on extraordinarily broad collections of books, newspapers, academic writing, historical documents, and countless other sources. A world where fair use is restricted would create several practical problems for users and institutions:
- Incomplete Answers: Users would receive less complete answers about current events, history, literature, and specialized subjects because the training data would be incomplete.
- Paywalled Knowledge: Greater dependence on paywalled material and multiple subscriptions would emerge, fragmenting access to information.
- Inconsistent Results: The same question might receive different answers depending on the user's employer, school, or subscription level.
- Reduced Coverage: Local, niche, minority-language, and out-of-print material would receive reduced coverage in AI systems.
- Educational Inequality: A widening gap would emerge between well-funded and underfunded educational institutions in their access to AI tools.
These consequences would extend far beyond the tech industry. They would affect every classroom, laboratory, startup, library, marketplace, and workplace that increasingly relies on generative AI.
What Do AI Developers Say About Fair Use?
Developers like OpenAI push back against the plaintiffs' interpretation of copyright law. They argue that their use of copyrighted material represents a clear case of fair use under decades of legal precedent. The material used in training isn't copied, stored, or resold; it's transformed into a new kind of tool. Additionally, developers cite a clear line of Supreme Court and appellate rulings suggesting that makers of multi-purpose products aren't usually liable when someone else uses that product to violate copyright law.
The irony, some observers note, is that authors and news outlets regularly rely on fair use themselves when writing their own books and articles. Yet many of these same entities are now going to court to stop AI developers from using copyrighted material in a similar transformative way.
What Is the Constitutional Argument for Fair Use in AI?
Supporters of fair use for AI training point to the Constitution itself. The Constitution states unambiguously that the purpose of copyright law is "to promote the progress of science and useful arts." Robust fair use has been and remains the principal way U.S. copyright law strikes the balance between protecting creators and enabling innovation.
The Constitution
Plaintiffs in the copyright cases argue that generative AI has taken us into unexplored and dangerous legal territory, comparing it to ancient maps that warned navigators to avoid uncharted waters by marking them with "Here there be dragons." But supporters of fair use argue that we already have a clear legal map in the form of established copyright doctrine and constitutional principles.
When Will This Legal Battle Reach a Decision?
After more than two years of preliminary legal jousting, the central issue of fair use is expected to come to a head in the fall, with arguments before a federal judge in New York City expected early next year. The outcome of this case could reshape not only the AI industry but also the future accessibility and affordability of artificial intelligence tools across society.