Why AI Copyright Settlements Are About to Get Much Bigger
The legal landscape for AI training data has fundamentally shifted. After years of uncertainty, courts are drawing clear lines between what's permissible and what isn't, and the financial consequences are staggering. Anthropic's $1.5 billion settlement in August 2025, the largest copyright recovery in U.S. history, reveals a critical distinction that will reshape how AI companies operate: using legally obtained copyrighted works to train models is likely fair use, but downloading from pirate sites is not.
What Changed in AI Copyright Law?
Two landmark rulings in June 2025 gave AI companies their first major courtroom wins, but with a crucial caveat. In Bartz v. Anthropic, Judge William Alsup of the Northern District of California ruled that using books to train Claude was fair use when those books were legally acquired, comparing it to a human reading widely to learn to write. A parallel ruling in Kadrey v. Meta reached the same conclusion for Meta's training practices. However, both cases drew a sharp distinction: legal acquisition is the threshold.
Judge Alsup ruled separately that Anthropic's downloading of pirated books from shadow libraries Library Genesis and Pirate Library Mirror was not protected by fair use. That finding drove Anthropic to settle the class action for $1.5 billion in August 2025, covering roughly 500,000 works at approximately $3,000 per title. Anthropic also agreed to destroy the original pirated files.
The settlement resolved past claims only. It does not license Anthropic's future training or cover outputs from its models. Claims filed after August 25, 2025, are explicitly excluded.
How Are Different Countries Approaching AI Training Data?
The U.S. is not alone in cracking down on unauthorized use. France's competition authority fined Google 250 million euros for using news articles without permission in training Gemini, signaling that European regulators are willing to act on unauthorized use of journalistic content in AI systems.
The UK High Court issued the first UK judgment directly addressing copyright infringement in the development of generative AI, in Getty Images v. Stability AI. The court rejected Getty's secondary copyright infringement claim, finding that Stable Diffusion's model weights did not constitute "infringing copies" under UK law. Getty did win a narrow trademark infringement finding on watermark reproduction from early model versions but was ordered to pay 69.4% of Stability's costs, making the victory financially pyrrhic.
The EU AI Act represents the most significant new regulatory development for AI and copyright globally. Under Article 53, all providers of general-purpose AI models, including foundation models like GPT, Claude, and Gemini, must publish a structured public summary of their training data and implement a policy complying with EU copyright law, including respecting opt-outs under the EU Copyright Directive's text and data mining exception. The European Commission published its mandatory template for training content disclosures on July 24, 2025.
Japan's approach remains the most permissive among major economies. Copyrighted works can generally be used for AI training provided the material itself is not from infringing sources, and the use does not unreasonably harm the copyright holder's interests.
Steps to Protect Your AI Training Practices
- Document Data Sourcing: Maintain detailed records of where training data originates, proving legal acquisition rather than pirated sources. This documentation is now critical evidence in copyright disputes.
- Implement Licensing Agreements: Explore licensing deals with content creators and publishers to ensure authorized use of copyrighted material in training datasets.
- Track Model Outputs: Monitor whether your AI models can reproduce substantial portions of protected works, as the U.S. Copyright Office's Part 3 report found that model weights themselves may infringe the reproduction right if they have memorized substantial protectable expression from training data.
Can AI-Generated Works Be Copyrighted?
The answer depends on jurisdiction, but the common thread is consistent: human authorship is required. The U.S. Copyright Office's January 2025 Part 2 report confirmed that AI outputs qualify for copyright protection only where humans provide sufficient creative input. The threshold is not minimal; writing a text prompt does not qualify.
In September 2022, the U.S. Copyright Office made history by issuing a groundbreaking registration for the comic book Zarya of the Dawn, created using the text-to-image AI tool Midjourney. The author clarified that the artwork was AI-assisted, not solely AI-generated. She structured the story, designed the page layouts, and made artistic decisions to arrange the elements alongside the AI-generated images.
However, an award-winning AI-generated print that won a competition at the Colorado State Fair was denied copyright protection. The creator expressed that he spent numerous weeks curating the perfect prompts and manually identifying the finished product, yet the work still failed to meet the human authorship threshold.
What Does This Mean for Musicians and Artists?
The philosophical foundations of intellectual property law offer valuable insight into how to conceptualize the repercussions of AI on artists engaging in creative work. Three primary approaches shape the debate: incentives-based utilitarian theory, which justifies IP protection as a way to encourage innovation and creation; Lockean theory, which grounds IP rights in the labor and effort creators invest in their work; and personality-based theory, which views IP as an extension of the creator's identity and autonomy.
Each approach suggests different paths forward for AI regulation and design. The incentives-based approach raises concerns that AI tools could undermine the financial incentives that motivate human creators. The Lockean approach questions whether AI-assisted creation involves sufficient human labor to warrant protection. The personality-based approach emphasizes how AI tools might affect artists' sense of agency and control over their creative identity.
Musicians have been particularly vocal about these concerns, especially following controversies over "deepfake" hit singles like the AI-assisted soundalike "Heart On My Sleeve" that appeared in 2023. A host of concerns have arisen for practicing musicians, songwriters, producers, and composers, many of them clustering around the perception of a shifting relationship between creator and creation.
The key takeaway for creators is clear: the legal landscape is stabilizing, but it demands vigilance. Companies and individuals relying on AI for content creation must meticulously document their creative process when using AI, highlighting the human decisions and modifications made to ensure a strong claim to authorship. The distinction between legal and pirated training data is now legally binding, and the financial penalties for getting it wrong are substantial.