Research

How Concentrated Are AI Citations? It Takes 1,357 Domains to Reach Half

The concentration curve of the Machine Relations Index across 22,179 cited domains. The most-cited domain holds 1.81% of citations and the top ten hold 7.41%; concentration is a property of the source class and of engine coverage, and inside the median segment thirteen cited answers reach the top ten.

Published Machine Relations Research
Index Analysis
TopicsAI SearchCitationsMeasurementSource Selection

In the September 18, 2026 release of the Machine Relations Index, the most-cited domain on the open web — reddit.com, cited in 2,012 of 15,782 observed answer runs — holds 1.81% of all citations recorded. The top ten domains together hold 7.41%. To account for half of every citation the Index observed, you have to count down to the 1,357th domain. To account for 80%, you count down to the 6,643rd.

That is a concentrated distribution. It is not the distribution the category talks about.

The working assumption behind most AI-visibility strategy is that answer engines collapse the web to a short list, and that the practical question is how to join a handful of dominant sources. Measured across six engines, 125 days and 22,179 cited domains, the head is real but thin, and the mass of citation sits in a very long middle. This matters for a publisher because it changes the question from can I displace Reddit to which of the 21,679 domains below the grading floor am I, and what would move me.

What was measured #

Every figure below comes from the public release at machinerelations.ai/data/machine-relations-index.json, contract machine_relations_index_public_view_v2.0, methodology mri_score_v2.0, generated September 18, 2026. The observation window runs May 10 to September 18, 2026 — 125 days, 15,782 observed answer runs across six engines (ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, Perplexity), 22,179 cited source domains.

One definition carries the whole piece. For each domain the Index records runs_cited: the number of observed answer runs in which at least one engine cited that domain. Summed across all 22,179 domains, that is 110,877 domain-run citations, and it is the denominator for every share below. It is not the same as the raw citation-event count, which is higher because a single answer run can cite one domain several times. Shares of domain-run citations answer "how often was this domain one of the sources", which is the publisher's question.

A domain's citation rate is its runs_cited divided by 15,782. A domain publishes a confidence grade only after clearing an evidence floor of 10 observed runs across 7 distinct run dates.

The concentration curve #

Top N domains Share of all domain-run citations
1 1.81%
3 4.07%
5 5.39%
10 7.41%
25 10.82%
50 14.19%
100 18.56%
250 26.44%
500 34.75%
1,000 44.95%
2,500 61.05%
5,000 74.60%

The Gini coefficient of the distribution is 0.663: unequal, but not winner-take-all. A winner-take-all citation layer would put the first ten domains well past a quarter of the total. These ten hold less than a thirteenth.

The ten domains at the top, with their source-role classification from the release:

Rank Domain Source role Runs cited Citation rate
1 reddit.com Community and social platform 2,012 12.75%
2 youtube.com Search or media platform 1,431 9.07%
3 linkedin.com Community and social platform 1,075 6.81%
4 medium.com Editorial publication 807 5.11%
5 forbes.com Editorial publication 649 4.11%
6 nih.gov Academic and government source 485 3.07%
7 gartner.com Analyst and consulting research 454 2.88%
8 techradar.com Editorial publication 448 2.84%
9 g2.com Market database 442 2.80%
10 arxiv.org Academic and government source 410 2.60%

Six distinct source-role classes appear in those ten slots. There is no single class that owns the head.

The median indexed domain is cited twice #

The other half of a distribution is the part nobody publishes. Across 22,179 domains:

Measure Value
Maximum runs cited 2,012
99th percentile 49
90th percentile 9
75th percentile 4
Median 2
25th percentile 1
Mean 5.0
  • 10,381 domains — 46.8% of the index — were cited in exactly one answer run across 125 days.
  • 14,053 (63.4%) were cited in two runs or fewer.
  • 18,166 (81.9%) were cited in five or fewer.
  • Only 2,165 domains (9.8%) reach ten or more observed runs.
  • Only 71 domains are cited in 100 runs or more.

The confidence grades follow directly. Of 22,179 domains, 14 carry grade A, 56 carry grade B, 430 carry grade C, and 21,679 are still collecting — below the evidence floor, with no publishable rate. Being in the index is not an achievement; almost every domain the engines touched is in it. Clearing the floor is the achievement, and 500 domains have.

Breadth is just as thin. 17,186 domains — 77.5% — appear in exactly one segment (one subject category paired with one question shape). The median domain appears in one. Only 710 domains appear in five or more. The maximum is 98.

Read together: the typical cited domain is cited once or twice, in one segment, and has not been observed often enough to earn a grade. It was a source for one answer to one kind of question, once.

Concentration is a property of the source class, not of AI answers #

This is the finding that changes what a publisher should do, and it is invisible in any single aggregate number. Split the 22,179 domains by their source role and measure concentration inside each class:

Source role Domains Domain-run citations Top 10's share of the class Median domain
Community and social platforms 27 4,314 97.3% 14
Academic and government 413 3,872 46.0% 2
Analyst and consulting research 364 3,381 36.3% 2
Market databases 688 6,066 29.4% 2
Editorial publications 1,222 13,309 25.7% 2
Vendor-owned sources 823 11,153 18.8% 3
Uncategorized sources 18,623 67,043 1.9% 2

The spread runs from 97.3% to 18.8% — a fivefold difference in how much the top of a class takes.

Community and social platforms are a near-monopoly. The class has only 27 domains in it, and ten of those take 97.3% of everything the class is cited for. For a new community property, the class is effectively closed: the seats are taken by Reddit, LinkedIn, Substack, dev.to and Facebook, and the tail behind them is negligible.

Vendor-owned sources are the flattest class measured. 823 domains share 11,153 citations, the top ten take under a fifth, and the median vendor domain is cited three times — the highest median of any class. A vendor page competes against a wide, shallow field rather than against an incumbent.

Editorial publications sit between: 1,222 domains, the largest classified citation pool at 13,309, and a top ten taking about a quarter. There is room in the class, but the median editorial domain is cited twice, which is the same as the median across the whole index. Being an editorial publication is not, by itself, an advantage.

The practical consequence: "how hard is it to get cited" has no single answer. It has seven answers, and which one applies to you is decided by what kind of source you are, not by how good the page is.

The scarcest position is not volume, it is engines #

There is a second concentration in this release, and it is steeper than anything in the volume curve. Split the 22,179 domains by how many of the six engines cited them at least once:

Engines citing the domain Domains Share of all domain-run citations
1 14,969 21.46%
2 3,502 13.33%
3 1,793 13.80%
4 963 15.10%
5 605 14.19%
6 347 22.12%

347 domains are cited by all six engines. They carry more citation than the 14,969 cited by exactly one — 24,530 domain-run citations against 23,793. One domain cited by all six carries as much as forty-three cited by one. The median all-six domain was cited in 41 runs against a whole-index median of 2.

That group is not the media list the phrase suggests. Of the 347: 153 are unclassified, 85 vendor-owned, 58 editorial publications, 21 market and company databases, 13 analyst and consulting sources, 11 academic and government sources, 4 community platforms, 1 search or media platform, 1 wire service.

The engines differ enormously in how wide they cast. Over this window the release records citations to 9,497 distinct domains from Perplexity, 7,441 from Google AI Mode, 7,079 from Gemini, 5,138 from Claude, 4,877 from ChatGPT and 2,279 from Google AI Overviews. A domain cited by one engine has cleared one retrieval system's filter; a domain cited by six has cleared six filters that share very little machinery, each governed by its own user agent under RFC 9309.

This is where effort converts. Moving from one engine to two is a larger structural change than doubling volume on a single engine, and it is invisible in any citation-rate number.

Inside a segment, the entry threshold is low #

The whole-index curve overstates the difficulty, because no publisher competes against 22,179 domains at once. They compete inside a segment: one subject category paired with one buyer question shape. 85 of the 151 measurable segments in this release are published. Across them:

Measure Median published segment
Distinct domains cited in the segment 231
Share held by the segment's top ten 24.1%
Share held by the segment's leader 4.3%
Cited runs needed to reach the top 100 2
Cited runs needed to reach the top 25 7
Cited runs needed to reach the top 10 13
Cited runs held by the segment leader 34

Thirteen cited answers puts a domain in the top ten of the median segment, and seven puts it in the top twenty-five. Restricting to the 78 buyer-question segments and excluding the seven longer-running news segments, the number falls to twelve.

Segment Domains cited Cited runs for top 10 Leader's cited runs
AI Visibility & GEO, best tools 351 17 30
Cybersecurity, best tools 213 13 44
Enterprise Software, how buyers choose 172 8 28

Segments vary far more than the median implies. The most open in the release is the legacy news topic, where 4,025 domains share citations and the top ten hold 8.3%. The most closed is Consumer Finance comparisons, where 170 domains share citations, the top ten hold 42.1%, and acorns.com leads at a 38.41% citation rate. 20,704 of the 22,179 indexed domains — 93.35% — appear in at least one published segment, and 16,270 appear in exactly one of them. Most cited domains are cited about one thing.

What this corrects #

Machine Relations published an analysis on June 14, 2026 — AI Citation Concentration: Why Market Databases Capture Disproportionate Share — which described the distribution as a power law and read market databases as capturing disproportionate share. That analysis was run under MRI Score v1.1 over 7,124 domains and 28,870 source events, and it is labelled on the page as superseded by v2.0 on July 5, 2026.

At v2.0, over three times as many domains, the shape reads differently in two ways. The head is weaker than "power law" suggests in ordinary use: 1.81% for the leading domain, 7.41% for the top ten. And market databases are not the most concentrated class — at a top-ten share of 29.4% they sit fourth of seven, behind community platforms, academic and government sources, and analyst research. The v1.1 reading was not wrong about market databases being strong in absolute terms; g2.com and crunchbase.com are rank 9 and rank 16 overall. It was wrong that their class is where concentration lives.

Correcting our own earlier published reading is part of what a versioned instrument is for. Every figure in this piece is tagged to release mri_score_v2.0 generated 2026-09-18, and a later release can overturn it the same way.

Where the independent research agrees #

An empirical study of source coverage across LLM-based and traditional search engines, analysing 55,936 queries over six LLM search engines and two traditional engines, reports that LLM-based engines cite domains with greater diversity than traditional search engines, and that 37% of domains are unique to the LLM-based engines (arXiv:2512.09483). That is an independent measurement, on a different population with a different method, pointing the same way as the curve above: the citation layer of AI answers is wider than the intuition that it collapses to a few sites.

Two adjacent lines of work frame why the tail behaves as it does. A measurement framework distinguishing citation selection from citation absorption across AI search platforms argues that appearing as a citation and actually shaping the answer are separate dependent variables (arXiv:2604.25707). And an observational audit of Google AI Overviews on "Your Money or Your Life" queries, using rank- and provenance-conditioned citation measures, finds citation behaviour that is not explained by retrieval rank alone (Kakimov et al., PMLR v318). Neither study measures the concentration curve directly; both support the reading that source selection is its own process rather than a re-ranking of search results.

Access policy is the one structural factor a publisher directly controls. The Robots Exclusion Protocol is standardised as RFC 9309; Google documents its crawlers and the Google-Extended token separately from Search inclusion (common crawlers, crawler overview); OpenAI documents its own bots and their user agents (OpenAI bots, catalogued independently at Dark Visitors); Cloudflare has made per-site blocking of AI crawlers a one-click control (Cloudflare). Google's own guidance on appearing in AI features is published in Search Central (succeeding in AI search), and Search Console documents where AI-surface performance appears in its reports (Search Console help). Google has described AI Mode as a distinct search experience rather than a variant of the results page (Google). On the demand side, survey work finds users click source links less often when an AI summary is present (Pew Research Center), which is why citation share and traffic share have to be measured separately. The broad web corpora these systems are built on are themselves public and enumerable (Common Crawl).

What a publisher can check #

Every one of the 22,179 domains in this release has a public page. A domain profile lists its overall citation rate, its confidence grade, its rank in the full universe, and every segment in which it has been observed — for example reddit.com. Category and question-shape leaderboards list the top domains per segment, such as AI infrastructure. The Index is the entry point, and the full release is machine-readable.

Domains that clear the evidence floor in a published segment also carry a citable badge stamped with this release id, served at machinerelations.ai/index/domains/<domain>/badge.svg and embeddable from the profile page.

Four checks are worth making against the numbers above. Which percentile of the curve your domain sits in, from the full-universe rank on its profile. Whether it clears the evidence floor, or is one of the 21,679 still collecting. Which source-role class it was assigned to, because the class decides which concentration figure applies to you. And how many of the six engines have cited it, because that number, not the rate, is what the steep part of this release's distribution is made of.

Limits #

This is a measurement of the Index's own observation window and question basket, not of the whole web. The Index runs a fixed basket of buying and research questions; a domain cited heavily for questions outside that basket will look thinner here than it is. Domains cited zero times in the window are not in the file at all, so the 22,179 is a population of cited domains, not of eligible ones — the true denominator of "sites that could have been cited" is larger and unknown.

Nothing here is causal. The curve describes what engines cited over 125 days. It does not say why, and it does not say that any intervention moves a domain along it. A domain's position at the next release is the only test of that.

Finally, concentration measured on domain-run citations is not the same as concentration measured on citation events, on answer position, or on whether the cited page actually supported the claim beside it. That last question is measured separately in Answer-Source Fidelity, and its results do not follow from this curve.


Release: mri_score_v2.0, contract machine_relations_index_public_view_v2.0, generated 2026-09-18. Window: 2026-05-10 to 2026-09-18, 125 days observed. Population: 22,179 cited domains, 15,782 observed answer runs, 110,877 domain-run citations. Evidence floor: 10 observed runs across 7 distinct run dates. See also the Machine Relations Index Report for September 18, 2026 for the same release read by segment leader and engine.