Each AI search engine runs its own retrieval pipeline, and those pipelines do not agree on which sources deserve a citation. Across 10,661 observed answer runs and 17,266 measured domains, the Machine Relations Index finds that the same domain can appear in one engine's citations and be absent from another's — even when both engines answer the same category of question.
The Six Engines Do Not Cite the Same Sources #
The MRI v2 measures citation rates for six engines: ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, and Perplexity. Among the top 100 most-cited domains, 71 are cited by all six engines. The remaining 29 show measurable gaps — an engine either never surfaces that domain or surfaces it so rarely it falls below the evidence floor of 10 observations across 7 distinct run dates.
This divergence is consistent with independent research. Frase's analysis of cross-engine citation overlap found that pairwise source overlap between engines ranges from 16% to 59% — even Gemini and Google AI Mode, both Google products, shared only 27% of cited sources on identical queries. QuickSEO's research across 680 million citations found 62% brand disagreement across ChatGPT, Google AI Mode, and AI Overviews, with only 11% domain overlap between ChatGPT and Perplexity results.
Reddit is the clearest example in the MRI data. It holds the highest overall citation rate in the index at 12.6% — meaning it appears in roughly one of every eight answer runs. But that rate is driven entirely by four engines: Gemini, Google AI Mode, Google AI Overviews, and Perplexity. ChatGPT and Claude do not cite Reddit at measurable rates. A brand relying on Reddit threads for AI visibility would appear in Google's AI products and Perplexity but remain invisible to ChatGPT and Claude users.
The pattern is not limited to Reddit. The MRI data shows several high-citation domains where individual engines diverge:
| Domain | Overall rate | Cited by | Not cited by |
|---|---|---|---|
| reddit.com | 12.6% | Gemini, Google AI Mode, AI Overviews, Perplexity | ChatGPT, Claude |
| yahoo.com | 2.6% | ChatGPT, Claude, Google AI Mode, AI Overviews, Perplexity | Gemini |
| substack.com | 2.0% | Claude, Gemini, Google AI Mode, AI Overviews, Perplexity | ChatGPT |
| techcrunch.com | 1.5% | ChatGPT, Claude, Google AI Mode, AI Overviews, Perplexity | Gemini |
| facebook.com | 1.5% | Google AI Mode, AI Overviews, Perplexity | ChatGPT, Gemini, Claude |
| marketsandmarkets.com | 1.2% | Claude, Gemini, Google AI Mode, AI Overviews, Perplexity | ChatGPT |
Source: Machine Relations Index v2, 82-day observation window ending 2026-08-04. Citation rates reflect the share of observed runs where the domain was cited. "Not cited by" means the domain did not meet the evidence floor for that engine.
Why Engines Diverge: Retrieval Pipeline Differences #
Each engine follows a retrieval-augmented generation pipeline, but the implementation details differ enough to produce different citation sets. Research tracking 134 URLs across multiple AI engines found that sources cited by one engine are frequently absent from another's responses to the same query.
Three structural factors explain most of the divergence:
Index composition. Google's engines (AI Mode, AI Overviews) draw from Google's web index, which heavily represents Reddit, social platforms, and news outlets. ChatGPT and Claude use different retrieval backends — ChatGPT uses Bing and its own search, Claude uses a separate retrieval pipeline — and neither appears to index Reddit content at the same depth. As Kakimov et al. (2026) demonstrated in their audit of Google AI Overviews, the composition of the underlying index directly shapes which sources appear in generated answers. Fahlout's citation research quantified the gap: citation rates vary up to 615x across platforms, with ChatGPT citing at 0.59% while other engines cite at rates an order of magnitude higher.
Source-trust scoring. Engines weight authority signals differently. Gracker AI's analysis of citation patterns across ChatGPT, Google AI Overviews, Claude, and Perplexity found that each engine applies distinct trust heuristics. Google's products favor sources already ranking in traditional search. Perplexity operates as a retrieval-first engine that crawls the web independently with PerplexityBot, blends its own index with real-time search results, and prioritizes pages based on relevance, domain trust, and content freshness. Claude appears to favor long-form, well-structured content from established editorial and academic sources. Notably, Fahlout found that traditional domain authority explains only r² = 0.05 of citation behavior across engines — the signal that drives Google rankings has limited predictive power for AI citations.
Content extractability. An engine can only cite a source it can parse. Indexable AI's research on how ChatGPT, Gemini, and Claude read content found that structured, clearly-authored pages with explicit entity information get cited more consistently across engines. The Stacc's analysis of the LLM citation pipeline found that 44.2% of citations come from the first 30% of a page's content, and that content with semantic HTML tables receives 2.5x more citations than equivalent content in paragraph form. API Serpent's research confirmed that content specificity matters: a page stating a precise statistic with attribution is far more likely to be cited than one offering a vague claim on the same topic.
Source Type Performance Across the Index #
The MRI v2 classifies every measured domain into a source role. Aggregate citation rates by role reveal which types of sources engines prefer in aggregate:
| Source role | Domains measured | Avg citation rate | Top domain | Top rate |
|---|---|---|---|---|
| Community and social | 27 | 1.11% | reddit.com | 12.6% |
| Search or media platform | 10 | 1.13% | youtube.com | 10.6% |
| Wire distribution | 9 | 0.22% | businesswire.com | 0.8% |
| Vendor-owned | 823 | 0.11% | ibm.com | 2.8% |
| Editorial publication | 1,073 | 0.10% | medium.com | 6.6% |
| Market database | 614 | 0.08% | g2.com | 3.3% |
| Academic and government | 358 | 0.08% | nih.gov | 3.5% |
| Analyst and consulting | 364 | 0.08% | gartner.com | 3.6% |
Source: Machine Relations Index v2, 17,266 domains classified. Rates are the mean citation rate per domain within each role. "Top domain" is the highest-rated domain in that role.
Community platforms and search/media platforms have the highest average citation rates per domain — driven by a small number of heavily-cited domains (Reddit, YouTube, LinkedIn). But the distribution is steep: most domains in every category cite below 0.1%.
For B2B brands, the practical implication is that vendor-owned domains average an 0.11% citation rate across the index. The top vendor domain (ibm.com) reaches 2.8%, while most vendor sites are rarely cited. The gap between vendor-owned and editorial/analyst sources matters for citation architecture — brands that rely solely on their own domain for AI visibility are competing against sources with structurally higher citation rates.
What This Means for AI Visibility Strategy #
The cross-engine divergence measured by the MRI has a direct consequence: optimizing for a single AI engine's citation behavior risks losing visibility on the others.
A source that passes ChatGPT's retrieval filters may not pass Gemini's. A domain that Google AI Overviews cites regularly may never appear in Claude's responses. The MRI tracks which of the six measured engines cite each domain — and among the top 100 domains, the median is all six engines, but 29 domains show gaps where at least one engine does not cite them at measurable rates.
Three patterns from the data inform a cross-engine approach:
-
High-authority editorial and analyst sources get cited most consistently across all six engines. Forbes, Gartner, NIH, and G2 all appear in citations from every measured engine. These sources have established trust signals that survive different retrieval pipelines.
-
Community and social platforms show the widest engine divergence. Reddit's 12.6% citation rate is the highest in the index, but it comes entirely from Google-adjacent engines and Perplexity. Substack and Y Combinator are cited by five engines but not ChatGPT. Social proof that works on some engines is invisible on others.
-
Cross-engine citation coverage correlates with confidence grade. Domains with A-confidence ratings (the most evidence behind their measured rates) are more likely to be cited across all six engines. Domains with B or C confidence are more likely to show engine gaps. Building citability on multiple engines simultaneously builds more stable visibility than concentrating on one.
How Machine Relations Measures Cross-Engine Citation Behavior #
The Machine Relations Index v2 measures citation rates daily across six AI answer engines. Each domain receives a citation rate — the share of observed answer runs where that domain appeared in citations — published only after it clears an evidence floor of at least 10 observations across at least 7 distinct run dates. Domains are graded into confidence tiers (A, B, C, or collecting) based on how much evidence stands behind their measured rates.
The current measurement window spans 82 days (2026-05-10 through 2026-08-04) with 90,259 source events across 10,661 answer runs. The index tracks 17,266 domains across 31 published strata (category-question type pairs that have cleared the evidence floor).
This cross-engine measurement is what makes the divergence patterns visible. Without consistent measurement across all six engines using the same methodology, the differences in citation behavior would be anecdotal rather than quantified.
FAQ #
Do all AI search engines cite the same sources? #
No. Among the top 100 most-cited domains in the Machine Relations Index, 29 show measurable citation gaps — they are cited by some engines but not others. Reddit is the most extreme example: cited at a 12.6% rate overall, but only by four of six measured engines. ChatGPT and Claude do not cite Reddit at measurable rates.
Which AI engine cites the most diverse set of sources? #
Based on MRI v2 per-engine citation data, Google AI Mode and Perplexity appear in the citation records of the widest range of domains. Google's AI products benefit from the depth of Google's web index. Perplexity combines its own crawled index with real-time search, giving it access to a broad range of sources including community platforms, vendor sites, and academic content.
Why does Reddit rank #1 in AI citations but not appear in ChatGPT? #
Reddit's high citation rate (12.6%) is driven by Google-adjacent engines (Gemini, AI Mode, AI Overviews) and Perplexity, which crawl and index Reddit content. ChatGPT and Claude use different retrieval backends that do not index Reddit content at the same depth. This reflects a structural difference in index composition, not a quality judgment about Reddit's content.
How should brands approach cross-engine AI visibility? #
Brands should verify their citation status across multiple engines rather than optimizing for one. The MRI measures how many of the six engines cite each domain. Domains cited by all six engines tend to have higher confidence grades, meaning their citation rates are backed by more evidence and are more stable over time.
Last updated: August 4, 2026. Data from Machine Relations Index v2, observation window 2026-05-10 through 2026-08-04.