Apple's 'Public Access' Defense Could Rewrite AI Copyright Rules
Apple is mounting a novel legal defense in a copyright lawsuit that could fundamentally change how courts interpret copyright and digital access laws for artificial intelligence development. Rather than relying on traditional fair use arguments, Apple contends that publicly accessible YouTube videos do not qualify as protected content under the Digital Millennium Copyright Act (DMCA), potentially establishing a precedent that affects every major AI company.
What Is Apple's "Access vs. Use" Defense?
The case involves allegations that Apple scraped millions of publicly available YouTube videos to train internal AI foundation models. Instead of arguing that the use was fair, Apple is making a more fundamental claim: the videos were never legally "protected" in the first place because they required no passwords, subscriptions, or restricted credentials to view.
This distinction matters enormously. Apple argues that simply because information is copyrighted does not mean viewing it constitutes unlawful access under the DMCA. The internet is filled with copyrighted material that anyone may lawfully view, including newspaper articles, blog posts, photographs, and public GitHub repositories. If Apple's argument succeeds, it could narrow future DMCA claims involving publicly accessible websites and reshape how courts think about web scraping for AI training.
How Are Courts Currently Handling AI Copyright Claims?
The broader landscape of AI copyright litigation is evolving rapidly, with courts establishing clearer principles about intellectual property and artificial intelligence. One of the most significant recent developments is a landmark 1.5 billion dollar class-action settlement between AI developer Anthropic and authors and publishers whose books were allegedly used without authorization to train Claude, Anthropic's large language model.
Judge William Alsup concluded that training AI models on copyrighted books constituted fair use under US copyright law. However, the court also held that Anthropic infringed copyright by creating and maintaining a central digital library containing more than seven million books, distinguishing the use of works for AI training from the unauthorized storage of copyrighted material.
In another significant case, the UK High Court largely ruled in favor of Stability AI in its dispute with Getty Images. Getty Images abandoned its primary copyright infringement claims after failing to establish that the training of Stable Diffusion occurred within UK jurisdiction. The court concluded that the Stable Diffusion model weights do not store or reproduce copies of the original photographs; instead, the model learns statistical patterns from training data rather than retaining the underlying copyrighted images.
However, the court did identify limited trademark infringement relating to early versions of Stable Diffusion, finding that some AI-generated images contained distorted Getty Images or iStock watermarks. The court described these watermark issues as historic and confined to earlier versions of the model, with no evidence that newer releases continued to generate infringing outputs.
What Role Does Copyright Management Information Play?
Beyond access controls, another hotly contested issue in AI litigation involves Copyright Management Information (CMI), which includes author names, copyright notices, ownership information, licensing terms, and attribution data. Many AI training pipelines clean enormous datasets before model training, stripping metadata, removing filenames, discarding attribution, and deleting copyright notices.
Whether those actions violate Section 1202 of the DMCA has become one of the most important questions in AI litigation. In one early DMCA-only AI lawsuit, publishers alleged OpenAI removed copyright management information from news articles during AI training. The Southern District of New York dismissed the complaint after concluding the plaintiffs lacked sufficient legal standing because they failed to demonstrate a concrete injury tied to the alleged removal of CMI.
However, in the New York Times v. OpenAI case, portions of the DMCA claims survived dismissal. Judge Sidney Stein allowed certain Section 1202 claims to proceed while dismissing others without prejudice, illustrating that DMCA claims are highly fact-dependent rather than categorically barred. Similarly, The Intercept's complaint included significantly more detailed factual allegations and examples, helping the claims survive the pleading stage.
How to Build Defensible AI Data Governance
Regardless of how Apple's motion is decided, companies developing AI systems should implement comprehensive compliance procedures to reduce litigation risk. Legal experts recommend establishing written policies and practices that address multiple dimensions of data handling:
- Licensing and Permissions: Obtain explicit licenses for copyrighted works, especially for publishers, stock photography, music, software documentation, and proprietary databases that carry higher litigation risk.
- Data Collection Documentation: Maintain clear records of web scraping practices, API permissions, and dataset sourcing to demonstrate lawful access and use during litigation.
- Metadata Preservation: Retain copyright notices, author names, licensing terms, and attribution information throughout the AI training pipeline rather than stripping this data during preprocessing.
- Content Classification: Differentiate between publicly viewable pages, password-protected content, subscriber-only materials, contractual restrictions, and technical access controls to ensure compliance with both copyright law and the DMCA.
- Takedown Procedures: Establish formal processes for responding to copyright takedown requests and removing flagged content from training datasets promptly.
Modern AI companies should establish written policies addressing data collection, web scraping, copyright compliance, licensing, dataset auditing, takedown procedures, and metadata preservation. These policies may become valuable evidence in future litigation, demonstrating that a company acted in good faith and with reasonable care.
Why Does Apple's Case Matter for the Entire AI Industry?
If Apple prevails, the decision could establish an important principle: public accessibility alone does not create DMCA liability. Every major AI developer relies, at least in part, upon publicly available internet information, including OpenAI, Anthropic, Google DeepMind, Meta, Amazon, Apple, Microsoft, and xAI.
Even if Apple succeeds on its DMCA arguments, copyright claims do not disappear. Instead, litigation shifts toward different questions about fair use, transformative purpose, and market harm. These issues remain pending across numerous AI cases and will likely define the next phase of AI copyright litigation.
The broader pattern emerging from recent court decisions suggests that successful AI development increasingly depends not only on technical innovation but also on legally defensible data governance. The Anthropic settlement demonstrates that AI developers may face substantial financial exposure where copyrighted works are acquired or stored without authorization, while the Getty Images judgment illustrates that successful copyright claims against AI developers may depend on jurisdictional and technical considerations.
As courts continue to examine AI training and fair use, the distinction between accessing publicly available information and unlawfully circumventing technological protections will likely become central to how intellectual property law evolves in the AI era. Apple's motion may therefore become one of the most closely watched copyright decisions of the AI era, with implications extending far beyond the company itself.