Logo
FrontierNews.ai

Court Rules Perplexity's Data Scraping Violates Copyright Law, Even When Content Is Publicly Visible

A federal district court has largely sided with Reddit in its lawsuit against Perplexity AI and data-scraping tool provider SerpApi, finding that the companies violated the Digital Millennium Copyright Act (DMCA) by circumventing technological protections designed to prevent automated access to Reddit content. The ruling marks a significant legal setback for AI answer engines that rely on scraped web data and could reshape how companies like Perplexity source information for their systems.

What Did the Court Find About Perplexity's Data Collection Methods?

The court rejected a central argument from Perplexity and SerpApi: that circumventing Google's SearchGuard technology is permissible because the same Reddit content remains accessible to human users. The judge found that SearchGuard, which uses JavaScript challenges and CAPTCHAs to block automated systems from accessing Google search results, qualifies as a legitimate technological access control measure under DMCA law.

According to the court's reasoning, SearchGuard functions like "a facial-recognition technology programmed to open the door of a home for residents but not for other visitors." The fact that humans can still read the content does not mean the access control measure is ineffective; rather, it simply differentiates between automated and human users. SerpApi's tool allegedly uses proxy servers, fake user-agent strings, and other techniques to mimic human behavior and bypass these protections, allowing Perplexity to extract Reddit snippets and feed them into its retrieval-augmented generation (RAG) database, which powers its AI responses.

How Does This Ruling Affect AI Answer Engines?

The decision has broad implications for how AI companies source training data and real-time information. Reddit operates a platform with more than 100 million unique users daily and hosts billions of posts and comments spanning nearly two decades, making it an invaluable resource for AI training. The court acknowledged this value explicitly, noting that Reddit's conversational data is "invaluable to AI companies" for training large language models.

Critically, the court found that Reddit had plausibly alleged lost profits because SerpApi's circumvention tool eliminated Perplexity's incentive to enter into a licensing agreement with Reddit, as other AI companies like Google and OpenAI have done. This suggests that courts may view scraping as a substitute for legitimate licensing arrangements, potentially forcing AI companies to negotiate commercial agreements rather than rely on unauthorized data access.

What Claims Did the Court Allow to Proceed?

The district court denied most of the defendants' motions to dismiss Reddit's claims, allowing the following to move forward:

  • DMCA Circumvention Claims: Reddit's allegations that SerpApi and Perplexity circumvented access control measures under Section 1201(a)(1)(A) of the DMCA survived the motion to dismiss.
  • Trafficking in Circumvention Technology: The court allowed Reddit's claims under Sections 1201(a)(2) and 1201(b), finding that SerpApi's tool was designed or marketed for circumvention purposes.
  • Civil Conspiracy Claim: Reddit's state-law civil conspiracy claim was permitted to advance, though the court dismissed related unfair competition and unjust enrichment claims as preempted by federal copyright law.

The court did grant one motion to dismiss: SerpApi's challenge to Reddit's Section 1201(b) trafficking claim. The judge held that SearchGuard is an "access control" measure under Section 1201(a), not a "rights protection measure" under Section 1201(b), because it blocks access to content but does not control what users do with content once obtained.

Why Does the Court's Reasoning Matter for Content Creators?

The ruling reinforces that content creators and platforms have standing to bring DMCA claims even if they are not the original copyright holders. The court rejected SerpApi's argument that only copyright owners may sue under the DMCA, instead holding that the statute's broad language covers "any person injured" and anyone "arguably within the zone of interests" protected by the law. This expands the potential plaintiffs in future cases and strengthens the legal position of platforms like Reddit that host user-generated content.

The court also found that Reddit plausibly alleged reputational harm based on the defendants' circumvention of Reddit's "deletion" feature, which allows users to remove sensitive posts about topics like mental health, fertility, addiction, and trauma. By scraping content before users can delete it, Perplexity and SerpApi undermine user privacy expectations and the platform's ability to protect sensitive discussions.

How Are AI Companies Currently Managing Data Access?

The legal pressure on scraping-based models is mounting as AI companies increasingly pursue licensing agreements. Reddit has already negotiated commercial arrangements with Google and OpenAI, granting them access to Reddit content in exchange for compensation and restrictions on use. The court's finding that Perplexity had no incentive to license Reddit data because it could access it for free through SerpApi suggests that future rulings may force AI companies to choose between licensing or facing legal liability.

Meanwhile, other AI platforms are exploring different approaches to data sourcing and automation. Some companies are building shared AI systems that integrate multiple scheduling tools and data sources, allowing teams to manage recurring tasks and information retrieval through centralized workflows. By August 2026, platforms including Claude, ChatGPT, and Perplexity had all introduced scheduled task capabilities, enabling users to automate data collection and processing without relying solely on web scraping.

Steps to Understand Your AI Tool's Data Sourcing Practices

If you use AI answer engines or are considering adopting them for your organization, consider these practical steps:

  • Review Licensing Agreements: Ask your AI vendor whether they have licensing agreements with major content platforms and publishers. Companies with legitimate licenses are less likely to face legal challenges and may offer more reliable, up-to-date information.
  • Evaluate Data Freshness and Attribution: Determine whether the AI tool clearly attributes information to its sources and whether it can access real-time data through authorized channels. Tools that rely on scraping may provide outdated or unattributed content.
  • Assess Legal Risk: If your organization uses AI tools for sensitive applications, consult with legal counsel about the potential liability associated with tools that may rely on circumventing technological protections or unauthorized data access.

The Reddit v. SerpApi ruling is still in its early stages, with the case moving forward to discovery and potentially trial. However, the court's decision to allow most of Reddit's claims to proceed signals that judges are taking data protection and licensing seriously, even as AI companies race to scale their models with vast amounts of internet data.