India's Top Court Says AI Training on News Articles Is Fair Game,For Now
India's Delhi High Court has sided with OpenAI in a landmark ruling on whether AI companies can use copyrighted news articles to train their language models, finding that the practice qualifies as "fair dealing" under Indian copyright law. The decision, handed down on July 24, 2026, marks the first major judicial ruling in India on AI training and copyright, and it signals a broad pro-AI path on the question of whether scraping and storing copyrighted content for machine learning constitutes infringement.
The case pitted ANI Media, one of India's largest news agencies, against OpenAI. ANI alleged that OpenAI scraped, stored, and used ANI's copyrighted news articles without permission to train the large language models (LLMs), which are AI systems trained on vast amounts of text data to predict and generate human-like responses, underlying ChatGPT. ANI sought an interim injunction to stop OpenAI from storing or using its works.
What Did the Court Actually Decide?
Justice Amit Bansal rejected ANI's request for an interim injunction, but the ruling is more nuanced than a simple win for OpenAI. The court addressed four core issues and found different outcomes on each:
- Training and Storage: The court held that OpenAI's storage of ANI's works to train ChatGPT qualifies as "fair dealing" under Section 52(1)(a) of India's Copyright Act, 1957, and therefore does not constitute infringement.
- Output and Reproduction: The court found no prima facie evidence that ChatGPT's responses reproduced ANI's articles, partly because the articles ANI cited were published after ChatGPT's training cutoff date.
- Jurisdiction: The court confirmed that Indian courts have jurisdiction over OpenAI, even though the company's servers are located in the United States, because OpenAI actively targets Indian subscribers and collects fees from them.
- Fair Dealing Defense: The court applied a two-step test to determine whether OpenAI's use qualifies as fair dealing and concluded that it does.
How Did the Court Interpret "Fair Dealing" for AI?
The court's reasoning on fair dealing is where the decision becomes particularly significant for the AI industry. The judge held that Section 52 of India's Copyright Act is not merely an exception to copyright protection but rather an integral part of the statute that defines user rights. This framing is important because it shifts the conversation from "what exceptions allow copying" to "what rights do users have." The court applied two key tests: a purpose test (which enumerated purpose under Section 52 does the use fall under) and a fairness test (is the dealing actually fair).
On the purpose test, the court held that OpenAI's use qualifies as "private or personal use, including research." The court reasoned that LLM training is a form of research, and because the training dataset is never made available to the public, the use is private. Critically, the court rejected the argument that commercial use disqualifies a company from invoking fair dealing. The judge noted that India's Parliament deliberately omitted the word "non-commercial" from Section 52(1)(a), even though it explicitly included that restriction in other subsections of the same provision. This means a for-profit company like OpenAI can still claim fair dealing protection.
On the fairness test, the court identified three factors that weighed in OpenAI's favor:
- Limited Use: OpenAI uses ANI's works only for training its LLMs; the training data is never communicated to the public in natural language or tokenized form.
- No Market Substitution: ChatGPT's functions, such as content creation and research assistance, are fundamentally different from ANI's business of news syndication. The court found no evidence that ANI lost subscribers or advertising revenue.
- Quantifiable Harm: The court noted that ANI itself had offered to license its content to OpenAI for USD 7.5 million, demonstrating that any loss is quantifiable and can be compensated financially rather than through an injunction.
What About the Output Claim?
ANI also argued that ChatGPT's responses sometimes reproduced ANI's articles verbatim or near-verbatim. The court rejected this claim at the interim stage for several reasons. First, all the ANI articles ANI cited as evidence were published in August or September 2024, after the training cutoff dates for both GPT-4 (April 2022) and GPT-4o (April 2024). Therefore, those specific articles could not have been part of the training data.
Second, the court analyzed the examples ANI provided and found that ChatGPT's responses were not substantially similar to ANI's articles when compared in their entirety. The court also confirmed that public availability of ANI's news content does not nullify its copyright, but it noted that copyright in news articles protects the specific expression, not the underlying facts. This means the threshold for establishing substantial similarity in news content is higher than for creative works like song lyrics.
What Does This Mean for the AI Industry?
The ruling creates a significant precedent in India, a country with over 400 million internet users and a growing AI sector. However, it is important to note that this is an interim order, not a final judgment. The case will proceed to trial, where ANI will have the opportunity to present additional evidence and arguments.
The decision also signals a shift in how courts may interpret fair dealing in the context of AI. By recognizing that commercial research can qualify as fair dealing, and by rejecting the notion that lawful access is a prerequisite for invoking the defense, the court has created more breathing room for AI developers to use publicly available copyrighted content for training purposes.
However, the ruling does leave open the possibility that evidence of memorization and verbatim reproduction could be established at trial. This means the substantive questions about whether AI systems can memorize and regurgitate copyrighted content remain unresolved.
How Are Other Tech Companies Responding to AI IP Disputes?
While the Delhi High Court was ruling on copyright, another major IP battle was brewing in the United States. Apple has filed a high-profile lawsuit against OpenAI alleging trade secret misappropriation, signaling a new wave of intellectual property litigation in the AI industry. According to IP litigators, Apple's suit "lays out a roadmap of future AI litigation," suggesting that trade secret disputes may become as common as copyright cases in the coming years.
"This industry is going to be cutthroat," noted Manavdas, a partner at IP firm MBHB, in reference to the emerging landscape of AI trade secret litigation.
Manavdas, Partner at MBHB
The contrast between the Delhi High Court's pro-AI stance on copyright and the emerging trade secret litigation in the United States highlights a key tension in AI regulation: different jurisdictions and different types of intellectual property claims may produce very different outcomes. Copyright law, at least in India, appears to favor broad fair dealing protections for AI training. Trade secret law, by contrast, may offer stronger protections to companies that claim their AI models or training methods are proprietary.
Steps for Companies Navigating AI Copyright and IP Risk
As the legal landscape around AI and intellectual property continues to evolve, companies operating in this space face a complex set of challenges. Here are key considerations for managing IP risk in AI development:
- Jurisdiction Matters: Understand that different countries have different copyright frameworks. India's fair dealing doctrine, as interpreted in ANI v. OpenAI, is more permissive than some other jurisdictions. Companies should assess their exposure in each market where they operate or collect data.
- Document Your Training Data Sources: Keep detailed records of where your training data comes from and how you accessed it. The court's finding that publicly available content without paywalls is fair game may not apply to all jurisdictions or all types of content.
- Monitor for Memorization: Even if training on copyrighted content is legal, outputting verbatim or near-verbatim reproductions of that content may not be. Companies should implement safeguards to detect and prevent memorization in their models.
- Prepare for Trade Secret Litigation: As Apple's suit against OpenAI demonstrates, trade secret claims may become as important as copyright claims in AI disputes. Protect your model architectures, training methods, and other proprietary information with the same rigor you would apply to trade secrets in other industries.
- Consider Licensing Agreements: The court noted that ANI's offer to license its content to OpenAI for USD 7.5 million shows that copyright holders may be willing to negotiate. Proactive licensing agreements can reduce legal risk and provide a revenue stream for content creators.
The ANI v. OpenAI decision is a watershed moment for AI and copyright law in India, but it is far from the final word. As the case proceeds to trial and as other jurisdictions grapple with similar questions, the legal landscape will continue to shift. For now, the ruling suggests that AI companies have more latitude to use publicly available copyrighted content for training purposes than many copyright holders had feared, at least in India.