Amazon and Meta's AI Bots Are Scraping the Web at Record Speed,Here's What It Means for Publishers
Amazon and Meta's artificial intelligence bots are responsible for the majority of automated content scraping across the internet, according to new data analysis of billions of website visits. The two companies' bots accounted for nearly two-thirds of all AI-related web traffic in the year ending May 29, 2026, raising fresh concerns about unauthorized content harvesting and the financial burden on publishers.
Data and technology company 51Degrees analyzed traffic patterns across three billion website visits from sites in telecommunications, advertising, media, retail, financial services, and technology sectors. The findings reveal a stark concentration of AI bot activity among just a handful of companies, with AmazonBot alone representing more than 4% of all web visits and Meta-External-Agent accounting for nearly 3%.
Which AI Bots Are Consuming the Most Website Traffic?
On May 29, 2026, AmazonBot generated 6.43 million website sessions, while Meta-External-Agent created 4.23 million sessions. Together, these two bots represented 63% of all sessions from the 72 AI-related bots identified by 51Degrees over the 12-month period. The next most active bots were AhrefsBot, which powers the Ahrefs marketing intelligence platform and Yep search engine, and Microsoft's BingBot, which crawls pages for Bing search results.
- AmazonBot: Generated 6.43 million sessions on May 29, representing more than 4% of total web visits and making broad, high-volume sweeps that fetch far more pages per visit than traditional search crawlers.
- Meta-External-Agent: Accounted for 4.23 million sessions on May 29, representing nearly 3% of total web visits and making it the second-largest AI bot by traffic volume.
- AhrefsBot: Generated 2.17 million sessions on May 29, representing 2.3% of total web sessions and powering both the Ahrefs platform and Yep search engine.
- Microsoft's BingBot: Created 1.62 million sessions on May 29, representing 1.3% of total web sessions and used to discover and index pages for Bing search results.
- Anthropic's ClaudeBot and OpenAI Search Bot: Ranked seventh and ninth respectively, generating approximately 288,000 and 129,000 sessions on May 29.
The concentration of traffic among these top bots underscores a growing imbalance in how AI companies access online content. AmazonBot's approach of making "broad, high-volume sweeps that fetch far more pages per visit than a search crawler" differs fundamentally from traditional search engine indexing, according to bot traffic analytics platform Known Agents.
Why Is AI Bot Traffic Growing So Rapidly?
The surge in AI bot activity reflects the broader race among technology companies to train large language models (LLMs), which are AI systems trained on vast amounts of text data to generate human-like responses. During the 12 months ending May 2026, websites monitored by 51Degrees experienced a 62.5% increase in bot traffic from AI companies. AI bots accounted for 13% of all web sessions in May 2026, up from 8% a year earlier.
This rapid growth has created a dual problem for publishers. Not only do these bots consume significant bandwidth and server resources, but they also extract content that AI companies then use to generate answers that compete directly with the original publishers' traffic. James Rosewell, CEO of 51Degrees, explained the stakes: "Whilst AI bots only represent a small percentage of overall bot traffic they are the most insidious because they not only consume vast quantities of bandwidth but they also use that stolen content to create AI answers that steal traffic from the original creator of that content".
James Rosewell, CEO of 51Degrees
"The increase in AI bot traffic to websites is a huge licensing revenue opportunity for publishers, but left unchecked is a growing and unwanted burden, particularly for smaller, independent players who can ill afford the impact of IP theft and increased bandwidth costs," said James Rosewell.
James Rosewell, CEO of 51Degrees
How Can Publishers Protect Their Content From AI Bots?
Publishers have several tools at their disposal to manage AI bot traffic, though adoption remains inconsistent across the industry. The most common approach involves using robots.txt, the Robots Exclusion Protocol, which is designed to tell automated software which parts of a website should not be accessed.
- Robots.txt Implementation: Publishers can add specific AI bot names to their robots.txt files to block unwanted scrapers, though only around 21% of publishers are currently blocking AI scraper bots this way, according to Known Agents.
- Comprehensive Bot Blocking: Publishers like Adweek, The San Diego Tribune, and the Arkansas Democrat Gazette have achieved 100% coverage of AI bots tracked by Known Agents, while others like W magazine, Screenrant, and CBS News have implemented blocking for only 2% of tracked bots.
- Bot Protection Services: 51Degrees has developed a bot protection and payment product that charges AI bots for unwanted data harvesting, offering publishers a potential revenue stream while controlling access to their content.
- Third-Party Scraping Workarounds: Publishers have raised concerns that AI companies can circumvent direct blocking by using third-party companies to scrape content instead, making comprehensive protection more difficult.
The good news for publishers is that AI scrapers and data providers generally respect robots.txt requests, with Known Agents finding that 95.6% of AI scrapers abide by these exclusion protocols. The primary issue is that publishers are not being comprehensive enough in adding bots to their robots.txt files.
However, while robots.txt provides a technical solution, it relies on voluntary compliance. Rosewell noted that "there is no technical solution that can stop the bots entirely," highlighting the need for both technical measures and policy-level solutions to address the growing challenge of AI content scraping.
Rosewell
The data from 51Degrees reveals a critical moment in the relationship between AI companies and content publishers. As AI bots continue to grow in volume and sophistication, publishers face mounting pressure to either monetize their content through licensing agreements or implement stronger protections. The current landscape suggests that without more aggressive action from publishers or regulatory intervention, AI companies will continue to harvest content at scale, fundamentally reshaping how online content flows through the AI training pipeline.