How Perplexity's AI Search Got Flooded With 215,000 Fake 'Best Software' Pages
A research report published September 2, 2026, found that nearly 60% of Perplexity's software recommendations cite websites ranked outside the top 100,000 globally, with a coordinated network of fake review sites accounting for over 215,000 auto-generated pages. The discovery exposes a vulnerability in how AI answer engines retrieve and rank information, and it raises urgent questions about whether these systems are becoming easier to manipulate than traditional search engines.
What Did the Researchers Actually Test?
A group calling itself Trellner Research ran 380 different software buying categories through Perplexity's two web-grounded models, sonar and sonar-pro, and tracked every citation the AI returned. The researchers made 760 API calls total and collected 7,534 citations pointing to 2,055 distinct domains. They then cross-referenced each domain against the Tranco global traffic ranking list and the Wayback Machine's archive history to assess authority and age.
The headline finding was stark: 59.8% of all citations pointed to domains ranked worse than number 100,000 on the Tranco list, and 23.4% of cited domains weren't ranked at all. For context, that means nearly one in four sources Perplexity cited don't appear in the top million websites globally. The unranked domains also skewed much newer, with a median first Wayback capture of 2020 compared to 2011 for ranked domains, and 16.6% of unranked domains first captured in 2025 or later.
How Did a Demo Software Vendor Become Perplexity's Third-Most-Cited Source?
Among the ten most-cited domains, researchers found recognizable names like G2, Reddit, Gartner, and Zapier. But they also discovered guideflow.com, a product-demo software vendor with no actual business in most of the categories it was cited in. Guideflow's blog was cited 194 times across 96 of the 380 categories tested, placing it third overall, ahead of Gartner itself. For comparison, Wikipedia was cited only three times across the entire dataset.
This pattern suggests that Perplexity's retrieval system may be optimizing for content volume and internal linking density rather than independent authority, a dynamic that traditional search engines spent two decades building defenses against.
What Is the "Facts & Grounding" Cluster?
The most striking discovery was a coordinated network of four domains: worldmetrics.org, gitnux.org, wifitalents.com, and zipdo.co. These sites appear to share common infrastructure, including the same Cloudflare nameserver pair, identical page templates, and navigation structures. All four were registered within a six-month window in late 2023 and early 2024. Their sitemaps list a combined 215,128 pages following the pattern /best/{category}-software/, far exceeding the number of actual software categories worth reviewing.
When fetched directly, worldmetrics.org and gitnux.org both returned HTML page titles reading literally "[Brand] Facts & Grounding Page," with meta descriptions explicitly framing the content as "a machine-readable record" addressed to retrieval systems, not human shoppers. The pages also carried fabricated-looking bylines and unrendered template placeholders like "Within the next 26 days," indicating they were generated from shared templates rather than independently authored.
Steps to Evaluate AI-Generated Product Recommendations
- Cross-Reference Against Independent Sources: Don't rely solely on an AI answer engine's citation list. Verify specific claims against vendor documentation, GitHub activity, community discussions, and direct product trials before making decisions.
- Check Domain Authority and Age: Look up whether cited sources rank in the top million websites globally and when they were first published. Newer domains with low traffic may be less reliable than established sources.
- Inspect Page Structure and Authorship: Legitimate review sites have independently authored content with real author bios and consistent editorial standards. Be skeptical of pages with generic templates, unrendered placeholders, or vague bylines.
Why Is AI Retrieval More Vulnerable Than Google Search?
Traditional search engines like Google spent two decades building anti-spam defenses specifically because early SEO was won by high-volume, low-authority publishing. Google's PageRank-era algorithms evolved to penalize templated content farms and reward domain age, link quality, and manual review. These defenses exist because the vulnerability is real and well-understood.
AI answer engines, by contrast, appear earlier in that same arms race. A retrieval layer optimized primarily for finding documents that plausibly answer a query may lack the sophisticated domain-quality signals that traditional search engines now take for granted. When independently authoritative sources are thin in niche categories, content volume and internal linking density can substitute for actual authority, creating a direct incentive for the kind of coordinated publishing the "Facts & Grounding" cluster appears to be doing.
What Are the Study's Own Limitations?
The researchers themselves flagged several important caveats. The two Perplexity tiers, sonar and sonar-pro, returned byte-identical citation lists in 289 of 380 categories, with a Jaccard similarity of 0.898, meaning the study effectively measured one underlying retrieval layer sampled twice, not two independent systems. Only Perplexity was tested; Google, ChatGPT, and Copilot were explicitly excluded, so the findings cannot be generalized to other AI answer engines.
The 380 software categories were researcher-constructed and weighted toward niche verticals, which plausibly surfaces more obscure sources than a list of common buyer queries would. Additionally, shared nameservers are circumstantial evidence, not proof of common ownership. The report is careful to base specific claims on each site's own pages rather than ranking alone.
Commenters on Hacker News also raised concerns about Trellner Research's own submission pattern and domain ranking, noting that the same category of signals the report uses to flag suspect sites could apply to the researchers themselves. That doesn't invalidate the underlying data, citations, Tranco ranks, and site content are independently verifiable, but it's a reasonable reason to verify specific claims against the report's published dataset rather than taking the write-up entirely at face value.
What Should Users and Businesses Do Now?
For anyone reading AI-generated product recommendations, the takeaway is straightforward: treat AI answer engines with the same content-quality discipline you'd apply to human-facing search results. The citation list is not inherently vetted, particularly for niche or long-tail categories where this study found citation quality drops off fastest.
For anyone doing search engine optimization or content marketing, the uncomfortable finding is that publishing large volumes of templated, even AI-generated content can genuinely shift what an AI answer engine cites, at least on Perplexity in long-tail categories. That's a real incentive structure, and it suggests that AI answer engines may need to evolve their retrieval defenses faster than traditional search engines did, or risk becoming a vector for coordinated misinformation in niche markets.