Why AI-Generated Art Can't Be Copyrighted, But Training Data Lawsuits Keep Piling Up
AI-generated content like text, images, and music cannot be copyrighted in the United States, according to the U.S. Copyright Office. However, the process of training AI models on copyrighted works remains in a legal gray area, with major lawsuits from The New York Times, Getty Images, and major newspapers challenging whether companies can use protected material without permission or compensation.
Can AI-Created Works Actually Be Copyrighted?
The short answer is no. The U.S. Copyright Office has long held that works created solely by artificial intelligence, even if prompted by a human, cannot receive copyright protection. This stance was reinforced through a landmark case involving "Zarya of the Dawn," a graphic novel created by Kristina Kashtanova using the Midjourney text-to-image generator.
When Kashtanova first registered the work in September 2022, the Copyright Office granted protection. But just months later, the office reconsidered and partially canceled the registration. The office determined that while Kashtanova's written text and the overall arrangement of elements could remain protected, the images themselves could not. The reasoning was straightforward: the images were "not the product of human authorship," but rather outputs generated by AI based on text prompts and training data.
Federal courts have backed this position. In August 2023, a judge in the U.S. District Court for the District of Columbia sided with the Copyright Office against computer scientist Stephen Thaler, who sought copyright protection for an image created by AI software.
What About Human-AI Collaboration?
The situation becomes more nuanced when humans and machines work together. According to intellectual property experts, the level of human control and creative contribution matters significantly. Daniel Gervais, a professor at Vanderbilt Law School who specializes in intellectual property law, explained the distinction.
"If a machine and a human work together, but you can separate what each of them has done, then copyright will only focus on the human part. It really needs to be an authorial kind of contribution. In that case, the fact that you worked with a machine would not exclude copyright protection," said Gervais.
Daniel Gervais, Professor at Vanderbilt Law School
The Copyright Office has since released updated guidance addressing all AI-human creative collaborations. The policy reiterates that if a human simply types a prompt and the machine generates complex written, visual, or musical works in response, the "traditional elements of authorship" have been executed by AI, a non-human entity. Therefore, such works receive no copyright protection.
Why Are Companies Suing Over Training Data?
While AI-generated outputs lack copyright protection, the data used to train these models is a different matter entirely. Many creators and companies argue that generative AI firms have scraped and used copyrighted material without permission or compensation. This has sparked a wave of lawsuits that could reshape how AI companies operate.
The legal disputes center on whether using copyrighted material to train AI models qualifies as "fair use," a doctrine that permits limited use of copyrighted material under certain conditions without needing the owner's permission. Pending lawsuits are directly challenging this assumption.
- Getty Images vs. Stability AI: Getty Images sued Stability AI, the company behind Stable Diffusion, for copying and processing millions of copyrighted images and their associated metadata without permission or compensation.
- The New York Times vs. OpenAI and Microsoft: The Times sued in 2023 for using millions of NYT articles to train AI models without compensation, and again in 2025 sued Perplexity for scraping content and providing information that directly competes with the publication's offerings.
- Newspaper Industry Action: Additional U.S. newspapers, including The New York Daily News and the Chicago Tribune, have sued Microsoft and OpenAI, creating mounting legal challenges for the AI companies.
- Voice Rights: TikTok settled a 2021 lawsuit with voice actress Bev Standing, who claimed the company used her voice without permission for its text-to-speech feature.
The stakes are substantial. According to NPR's sources, if courts find that OpenAI illegally used Times articles to train its models, the company could be forced to destroy its large language model (LLM) dataset, which is the foundation of its AI systems, and rebuild it from scratch.
How to Understand the Legal Gray Area Around AI Training
- Fair Use Doctrine: Currently, U.S. law permits the use of copyrighted material under certain conditions without permission, but lawsuits are challenging whether AI training qualifies as fair use.
- Pattern Recognition Process: Generative AI systems work by identifying and replicating patterns in data, meaning they must learn from real human-created works to generate outputs in a particular style or format.
- Pending Legal Outcomes: Ongoing court cases may fundamentally change how copyright law applies to AI, potentially restricting how companies can train models or requiring them to license content from creators.
- Authorship Question: Courts and regulators are grappling with whether AI systems can be considered authors or whether they are merely tools that replicate patterns from human-created works.
The fundamental tension is this: generative AI systems cannot legally create copyrightable works, yet they were trained on copyrighted works that their creators did not license. As courts work through dozens of pending cases, the resolution could reshape the entire AI industry's approach to data sourcing and model training.