Why Public Broadcasters Blocking AI Search Engines May Backfire on Their Own Mission
Public broadcasters worldwide are using robots.txt files to block AI search engines and training systems from accessing their journalism, but this defensive move may contradict their fundamental mission to ensure universal access to reliable information. When AI systems cannot read content, growing numbers of people asking questions to AI tools won't see that journalism, effectively taking publicly funded news "off air" in the age of generative discovery.
How Are Public Media Outlets Currently Blocking AI Access?
A robots.txt file is a plain text instruction file that sits on every website and tells search bots and AI crawlers what content they can or cannot read. It's a simple but powerful tool, and it's publicly visible to anyone who adds "/robots.txt" to a domain name.
When researchers examined the robots.txt files of seven major public broadcasters, they found wildly inconsistent approaches. The Australian Broadcasting Corporation (ABC) and Special Broadcasting Service (SBS) had radically different block-lists with no obvious coherent rationale; one blocked 17 AI crawlers while the other blocked just 2. The BBC took the most aggressive stance, preventing AI systems from using its content for training, search, creating news summaries, or business use. The BBC even threatened legal action against AI startup Perplexity for ignoring its robots.txt and using its content without permission.
Some entries in these files read like time capsules, referring to long-dead content platforms and outdated commercial advertising arrangements. The inconsistency suggests these decisions were not made strategically but rather inherited from past website configurations or left to content management system defaults.
What Does This Mean for Public Access to News?
The consequences are measurable and troubling for public broadcasters' stated mission. Research by the Institute for Public Policy Research found that ChatGPT sourced more information from GB News, Al Jazeera, and Marie Claire than from the BBC, the UK's most popular and trusted news outlet. The BBC's absence from ChatGPT answers came at a direct cost to the public: when people ask AI tools questions, they're not seeing the journalism from the outlet most people trust.
The same problem appears in Google's Gemini, which generates the "AI Overview" in Google Search results. Because Google Search is ubiquitous, it's probably the most widely used AI answer engine even if many users don't realize they're using AI at all. When public broadcasters block these systems, their journalism simply disappears from the discovery environment where millions of people now seek answers.
How Should Public Media Outlets Balance Copyright and Access?
The Public Media Alliance defines the purpose of public media as providing "a variety of quality content that is universally accessible to a diverse audience" and enabling "citizens to access and interact with free, independent, engaging and relevant content whether they are in rural or urban environments, irrespective of economic status or technology." That last phrase, "irrespective of technology," is crucial in an era when people increasingly get their news by asking questions to AI systems.
The tension is real: public broadcasters have legitimate concerns about copyright, training data, and how their content is used. But the BBC's approach, which blocks search and AI access entirely, may be defending intellectual property at the expense of the public mission. Even if the BBC didn't want its content used for training AI models, it could still allow its journalism to be searched and cited by AI systems. Instead, it chose a blanket ban.
The strategic decision about what access to give search and AI bots cannot be left to technical defaults or whoever manages the website. For public media, it's difficult to defend blocking search and AI if they are to fulfill their purpose of making content accessible regardless of technology. In an age when people reach for search and AI for answers to questions, blocking these systems is tantamount to taking yourself off air.
Steps Public Broadcasters Can Take to Maintain Visibility in AI Discovery
- Audit robots.txt files strategically: Review what bots are currently blocked and make deliberate decisions rather than inheriting outdated configurations. Distinguish between training bots (which may warrant blocking) and search and AI citation bots (which serve the public mission).
- Allow AI search and citation while protecting training: Permit systems like Perplexity, ChatGPT, and Google's AI Overviews to search and cite journalism, even if you block systems used specifically for training large language models. This preserves visibility without sacrificing copyright concerns.
- Monitor AI visibility alongside traditional rankings: Track whether your journalism appears in AI-generated answers and comparisons, not just traditional search rankings. Visibility in AI systems is now as important as appearing on a search results page.
- Maintain clear, structured content: Write journalism with question-based headings, clearly defined concepts, and regular updates. Content that's easy for humans to understand is also easier for AI systems to accurately cite and reference.
The challenge facing public broadcasters reflects a broader shift in how people discover information. A generation ago, most people got news from newspapers, radio, and television. Today, they increasingly get it from a device by asking questions answered by AI. Content that an AI system cannot read is content a growing share of the public won't see. For institutions founded on the principle of universal access, that's a mission problem, not just a technology problem.