AI search engines cite brands that are retrievable, extractable, and corroborated by independent sources. Machine Relations Index data covering 17,342 domains and 90,971 citation events across six engines — ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, and Perplexity — shows that source type, entity clarity, and cross-engine presence matter more than domain authority or organic rank.
What the citation rate data shows #
The Machine Relations Index v2 measures citation rates: how often each domain appears in AI-generated answers across observed runs. A domain's citation rate is published only after it clears the evidence floor of at least 10 observations across at least 7 distinct run dates.
Out of 17,342 measured domains, 348 have earned enough evidence for a confidence-graded citation rate. The remaining 16,994 are still collecting.
The top 10 most-cited domains span five source roles:
| Rank | Domain | Citation Rate | Source Role | Confidence | Engines |
|---|---|---|---|---|---|
| 1 | reddit.com | 12.66% | Community / social | A | 4 |
| 2 | youtube.com | 10.63% | Platform / search | A | 6 |
| 3 | linkedin.com | 7.57% | Community / social | A | 6 |
| 4 | medium.com | 6.63% | Editorial media | A | 6 |
| 5 | forbes.com | 4.30% | Editorial media | A | 6 |
| 6 | gartner.com | 3.61% | Analyst research | A | 6 |
| 7 | nih.gov | 3.48% | Academic / government | A | 6 |
| 8 | g2.com | 3.30% | Market database | A | 6 |
| 9 | arxiv.org | 3.22% | Academic / government | A | 6 |
| 10 | ibm.com | 2.77% | Vendor-owned | B | 6 |
Source: Machine Relations Index v2, 83-day window (2026-05-10 to 2026-08-05), 10,733 observed runs.
This concentration at the top matches broader industry findings. Hexagon's analysis of 100,000 AI citations found that the top 2% of e-commerce brands capture 78% of all AI search recommendations across ChatGPT, Perplexity, Claude, and Google AI Overviews — a concentration that exceeds even paid search dominance. Searchless.ai's entity study found that only 2.4% of 500 tracked brands achieved consistent cross-engine citations, with 42% not mentioned by ChatGPT at all when asked for category recommendations.
Source role determines citation ceiling #
Individual domain performance matters, but source role sets the structural ceiling. Averaging citation rates across all domains with at least 10 cited runs reveals a clear hierarchy:
| Source Role | Avg Citation Rate | Domains Measured | Total Citations |
|---|---|---|---|
| Platform / search | 2.83% | 4 | 1,214 |
| Community / social | 2.12% | 14 | 3,180 |
| Wire distribution | 0.58% | 3 | 188 |
| Market database | 0.40% | 95 | 4,126 |
| Academic / government | 0.39% | 57 | 2,416 |
| Vendor-owned | 0.38% | 196 | 7,996 |
| Editorial media | 0.36% | 226 | 8,794 |
| Analyst research | 0.30% | 71 | 2,277 |
Community and platform sources earn citation rates 5–7x higher than editorial or analyst sources. This does not mean editorial sources are unimportant — editorial media accounts for the most total citations (8,794) across the most measured domains (226). The rate difference reflects how engines weight different source types for different question shapes.
DigitalApplied's ranking factors study confirms the content-type breakdown: blog and editorial content pages account for 53.46% of all AI citations, followed by news at 14.09%. Press releases via wire services account for just 0.04% — consistent with the MRI finding that wire distribution earns the lowest citation rates across all source roles.
Market databases like G2 (3.30% citation rate, confidence A) and Crunchbase (2.52%, confidence B) outperform their role average because they provide structured, extractable comparison data that AI engines use to ground product-category answers.
What separates cited brands from mentioned ones #
Being mentioned in an AI answer is not the same as being cited. Semrush's ghost citations study quantified the gap: 62% of all AI citations are "ghost citations" where the source is attributed but the brand itself is never named in the answer text. The study found that 74.9% of brand appearances in AI answers included citations, while only 38.3% included brand mentions — making the citation rate nearly double the mention rate.
The gap varies by engine. Semrush found that Gemini shows an 83.7% mention rate but only a 21.4% citation rate, while ChatGPT operates in reverse: 87% citation rate but just 20.7% mention rate. This means a brand can be a cited source on ChatGPT without ever being named in the answer, or mentioned frequently by Gemini without receiving source attribution.
Research from Zatuchin et al. (2026) analyzed 167,551 URL-grounded citations across 128 brands in 12 markets and found that LLMs ground brand reputation answers in retrieved web sources — and those sources determine what the model says. The brand that controls the best-structured, most-retrievable source about itself earns the citation.
Five factors that drive citation selection #
Cross-referencing the MRI data with external research identifies five consistent citation selection factors.
Entity clarity. AI engines match entities before they match content. DeepSmith's analysis of how LLMs recognize and match entities shows that brands with clear entity signals — consistent naming, structured data, and cross-platform presence — are recognized as citable sources. Hexagon's data quantifies the impact: brands with Wikipedia or Wikidata presence are cited by Perplexity at 4.7x the rate of those without it.
Third-party corroboration. Third-party editorial mentions from high-authority publications are 3.2x more predictive of AI citations than on-site content volume. DigitalApplied's study corroborates this with correlation data: branded web mentions correlate at 0.664 with AI visibility, compared to 0.218 for backlinks — roughly a 3x difference. Top-quartile brands average 169 AI Overview mentions versus 14 for the next tier.
Source architecture. The pages that earn citations share structural traits: answer-first formatting, explicit evidence, and direct definitions. Sprinklr's controlled study across six AI engines found that tightly focused pages answering single questions were cited far more often than broad "ultimate guides," and that substantive evidence and trust signals drove citations while formatting changes produced negligible results. A Moz study of approximately 40,000 queries found that 88% of Google AI Mode citations come from pages outside the traditional organic top 10, confirming that retrieval-stage source selection operates on different signals than organic ranking.
Statistic density. Getpromptive's analysis of 10,000 queries found that adding quantified data points increased AI citation visibility by 41% — the most actionable single signal in their study. Pages with specific numbers, percentages, and comparative data earn citations at higher rates than pages with qualitative assertions alone.
Content freshness. DigitalApplied's data shows AI-cited content is 25.7% fresher than organic search results: cited pages average roughly 2.9 years old versus 3.9 years for organic top-10 content. Perplexity is the most freshness-sensitive engine, with Reddit comprising 46.4% of its social citations, partly because Reddit surfaces recent discussions. Erlin.ai's study quantified the impact: content updated within 90 days has a 67% higher citation rate than older content, and AI engines cite an average of 2.8 brands per query — making freshness a direct lever for earning one of those limited slots.
Cross-engine presence as a citation predictor #
The number of engines citing a domain correlates with citation rate more reliably than any single content signal. Among the nine confidence-A domains in the MRI:
- All nine are cited by at least four engines
- Seven of nine are cited by all six engines
- Reddit, cited by only four engines, still leads at 12.66% — suggesting that citation intensity within individual engines can compensate for breadth
Getpromptive's research underscores this pattern from the opposite direction: only 11% of cited domains overlap between platforms. Each engine maintains a largely independent retrieval pipeline, so a brand visible on one engine may be invisible on another. The practical implication is that optimizing for a single engine leaves citation rate structurally capped.
This shifts where citations originate. The traditional organic top 10 is no longer the primary citation source: DigitalApplied's data shows only 38% of AI citations now come from Google's top 10 organic results, down from 76% in mid-2025. The remaining 62% come from pages ranking 11–100 (31.2%) or beyond rank 100 (31.0%).
The role top domains by category #
For brands evaluating where they stand relative to peers, source role context matters. The top-cited domain in each role:
| Source Role | Top Domain | Citation Rate | Confidence |
|---|---|---|---|
| Community / social | reddit.com | 12.66% | A |
| Platform / search | youtube.com | 10.63% | A |
| Editorial media | medium.com | 6.63% | A |
| Analyst research | gartner.com | 3.61% | A |
| Academic / government | nih.gov | 3.48% | A |
| Market database | g2.com | 3.30% | A |
| Vendor-owned | ibm.com | 2.77% | B |
| Wire distribution | businesswire.com | 0.79% | C |
Wire distribution earns the lowest citation rates, with the top domain at 0.79%. DigitalApplied's content-type data corroborates this: press releases via wire services account for just 0.04% of all AI citations.
AI answers also compress the competitive space. Sprinklr's research found that AI answers typically cite only three or four sources, compared to ten ranking positions in traditional organic search. Fewer available slots makes each citation more valuable and means brands compete against fewer but stronger peers for each answer.
What this means for Machine Relations #
Machine Relations — the discipline of managing how AI systems discover, evaluate, and cite a brand — requires understanding these citation selection mechanics before optimizing for them.
The combined MRI and external research data points to a structural conclusion: citation selection is a source architecture problem, not a content volume problem. Engines retrieve candidates from their crawl index, evaluate each candidate's extractability and trustworthiness, then cite the source that best grounds the answer. Brands that treat this as a content production challenge miss the structural layer.
The practical sequence for brands that want to earn AI citations:
- Verify entity recognition. Confirm your brand is recognized as a distinct entity by each engine. Brands with Wikipedia or Wikidata presence are cited at 4.7x the rate of those without it.
- Build third-party corroboration. Earned mentions from independent sources correlate at 0.664 with AI visibility — roughly 3x stronger than backlinks.
- Structure for extraction. Pages that lead with direct answers and include specific data points earn 41% higher citation visibility than qualitative pages.
- Measure across engines. With only 11% domain overlap between platforms, single-engine monitoring misses most of the picture. Citation rate — how often a source appears in AI-generated answers — is the relevant measurement.
FAQ #
Do AI search engines use the same ranking factors as Google organic search? No. While some overlap exists, AI engines select sources through retrieval-augmented generation, which evaluates extractability, answer relevance, and source trustworthiness differently. DigitalApplied's data shows only 38% of AI citations now come from Google's organic top 10, down from 76% in mid-2025. Branded web mentions are 3x more predictive of AI citation than backlinks.
Why does Reddit have the highest citation rate despite being on only four engines? Reddit's citation rate (12.66%) reflects its concentration in the engines that cite it — Gemini, Google AI Mode, Google AI Overviews, and Perplexity. Perplexity is especially reliant on Reddit, where Reddit comprises 46.4% of its social citations. Reddit's absence from ChatGPT and Claude's citation lists does not reduce its rate in the engines that use it.
What is the difference between an AI citation and an AI mention? A citation is a source-attributed reference with a link. A mention is when a brand name appears without attribution. Semrush's ghost citations study found that 62% of AI citations are ghost citations — the source is attributed but the brand is never named in the answer. Citation behavior also varies by engine: ChatGPT has an 87% citation rate with 20.7% mention rate, while Gemini shows 83.7% mention rate with only 21.4% citation rate.
How many domains does the Machine Relations Index measure? The MRI v2 measures 17,342 domains across 90,971 citation events over an 83-day window. Of these, 348 domains have accumulated enough evidence for a confidence-graded citation rate. The remaining domains are still collecting data. The evidence floor requires at least 10 observations across at least 7 distinct run dates.
Does domain authority predict AI citation rates? The data does not support a strong relationship. Hexagon's analysis found that third-party editorial mentions are 3.2x more predictive of AI citations than on-site content volume. Source role — whether a domain is a community platform, editorial outlet, or market database — is a stronger predictor of citation ceiling than domain authority.
Last updated: August 5, 2026. Data: Machine Relations Index v2, 83-day observation window. External research from Hexagon, Semrush, Sprinklr, DigitalApplied, Getpromptive, Zatuchin et al., DeepSmith, Searchless.ai, Erlin.ai, and ZipTie/Moz cited inline with source links.