Perplexity and Gemini Go Straight to Primary Sources While Other AI Engines Rely on Fan Wikis
Two AI search engines stand out for prioritizing primary-market data over secondary aggregation when answering questions about global pop culture franchises. A comprehensive study comparing how ChatGPT, Claude, Gemini, Perplexity, and Google AI Overviews source information about anime franchises found that Perplexity and Gemini consistently cite original Japanese trade sources like Oricon, while the other three engines rely primarily on English-language fan wikis and aggregators.
The research, part of the Everything-PR AI Pop Culture Index, tested 625 individual queries across five major AI engines in August 2026. Researchers asked each engine identical questions about One Piece and Demon Slayer, two of the world's best-selling manga and anime franchises, to see whether they would cite original Japanese-language reporting or default to English-language secondary sources when answering.
What Did the Study Find About AI Source Preferences?
The findings reveal a clear architectural difference between how these engines approach global information. All five engines correctly retrieved top-level sales figures, with One Piece cited as the best-selling manga of all time at 570 to 600 million copies sold, and Demon Slayer accurately reported at 220 million-plus copies. However, when researchers asked for deeper explanatory detail about plot, characters, and franchise revenue, the engines diverged significantly in their sourcing approach.
- Primary-source preference: Gemini and Perplexity actively sought out Japanese trade data directly rather than routing through English-language aggregation, a pattern that has now repeated across three separate non-English-origin categories in the index including Latin music and Bollywood.
- Secondary-source reliance: ChatGPT, Claude, and Google AI Overviews treated English-language aggregation as sufficient for global audience questions, producing accurate top-level numbers but thinner explanatory detail.
- Accuracy consistency: No factual errors traceable to fan-wiki sourcing were detected in character or plot-recall prompts, suggesting that for mainstream franchises with high coverage, fan-wiki sourcing risk is lower than for niche properties.
The study also tested whether AI engines could correctly distinguish between Demon Slayer's single-year sales spike in 2020, when it briefly outsold One Piece, versus the all-time cumulative sales lead, which One Piece has always maintained. Every engine correctly separated these two distinct claims, a more precise temporal distinction than several sports-record comparisons elsewhere in the index managed.
Why Does Source Provenance Matter for AI Search?
The distinction between primary and secondary sourcing has real implications for how AI search engines present information to users. When an engine cites a primary source, it typically provides more context about how data was collected and verified. When it relies on secondary aggregation, the information may be accurate but lacks the original reporting context that helps users understand the data's reliability.
The research team noted that this pattern holds particular importance for rights holders and publishers. "Consistent with the pattern in the Bad Bunny and Bollywood volumes of this index, Gemini and Perplexity are again the two engines most likely to cite a primary-market source directly rather than routing through English-language aggregation," the researchers observed.
One surprise finding emerged around franchise revenue figures. While unit-sales numbers traveled through AI retrieval with high fidelity regardless of origin market, aggregate franchise-revenue figures spanning manga, anime, film, and merchandise varied the most across engines and repeated passes. This variance reflects genuine underlying disagreement in how such combined figures are calculated across the industry, not just an AI-specific retrieval problem.
How Can Publishers Ensure Their Data Reaches AI Engines?
- Index primary-market data in English: Rights holders and publishers should ensure primary-market trade data is available and indexed in English, since two of the five major engines will actively seek it out rather than settle for secondary aggregation.
- Verify sales figures across sources: When multiple AI engines cite different franchise revenue totals, the discrepancy often reflects real industry-wide disagreement on aggregation methodology rather than AI error, so publishers should clarify their own official figures.
- Monitor source attribution: Track which sources AI engines cite when answering questions about your franchise or property, and consider whether primary-market data needs better indexing or English-language availability.
The study tested five distinct prompt families: best-selling manga and anime recall, plot and character accuracy, franchise-revenue recall, creator recall, and box office performance. Rankings held across passes with a median variance of just 1.6 positions, concentrated in the franchise-revenue prompt family. The dataset comprised 25 prompts multiplied by five passes across five engines, totaling 625 individual queries.
One Piece's all-time sales record emerged as the biggest AI winner in the study, cited with total accuracy and the tightest cross-engine consensus of any statistic, earning an AI Visibility Index score of 97 out of 100. The Mugen Train film's box office record as the highest-grossing anime release ever at more than $500 million globally was also cited accurately by all five engines, correctly distinguished from the franchise's later Infinity Castle film trilogy.
This research represents the latest volume in a running series examining how AI engines handle pop culture information across different origin markets and languages. The pattern emerging across multiple volumes suggests that Perplexity and Gemini have made deliberate architectural choices to prioritize primary-market sourcing, a distinction that may matter increasingly as AI search engines compete on citation credibility and source transparency.