# How Concentrated Are AI Citations? It Takes 1,357 Domains to Reach Half

The concentration curve of the Machine Relations Index across 22,179 cited domains. The most-cited domain holds 1.81% of citations and the top ten hold 7.41%; concentration is a property of the source class and of engine coverage, and inside the median segment thirteen cited answers reach the top ten.

Canonical URL: https://machinerelations.ai/research/how-concentrated-are-ai-citations-source-distribution-2026
Published: 2026-09-18
Research type: Index Analysis
Tags: ai-search, citations, measurement, source-selection

## Source Body

In the September 18, 2026 release of the Machine Relations Index, the most-cited domain on the open web — reddit.com, cited in 2,012 of 15,782 observed answer runs — holds 1.81% of all citations recorded. The top ten domains together hold 7.41%. To account for half of every citation the Index observed, you have to count down to the 1,357th domain. To account for 80%, you count down to the 6,643rd.

That is a concentrated distribution. It is not the distribution the category talks about.

The working assumption behind most AI-visibility strategy is that answer engines collapse the web to a short list, and that the practical question is how to join a handful of dominant sources. Measured across six engines, 125 days and 22,179 cited domains, the head is real but thin, and the mass of citation sits in a very long middle. This matters for a publisher because it changes the question from *can I displace Reddit* to *which of the 21,679 domains below the grading floor am I, and what would move me*.

## What was measured

Every figure below comes from the public release at [machinerelations.ai/data/machine-relations-index.json](https://machinerelations.ai/data/machine-relations-index.json), contract `machine_relations_index_public_view_v2.0`, methodology `mri_score_v2.0`, generated September 18, 2026. The observation window runs May 10 to September 18, 2026 — 125 days, 15,782 observed answer runs across six engines (ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, Perplexity), 22,179 cited source domains.

One definition carries the whole piece. For each domain the Index records `runs_cited`: the number of observed answer runs in which at least one engine cited that domain. Summed across all 22,179 domains, that is **110,877 domain-run citations**, and it is the denominator for every share below. It is not the same as the raw citation-event count, which is higher because a single answer run can cite one domain several times. Shares of domain-run citations answer "how often was this domain one of the sources", which is the publisher's question.

A domain's citation rate is its `runs_cited` divided by 15,782. A domain publishes a confidence grade only after clearing an evidence floor of 10 observed runs across 7 distinct run dates.

## The concentration curve

| Top N domains | Share of all domain-run citations |
|---|---|
| 1 | 1.81% |
| 3 | 4.07% |
| 5 | 5.39% |
| 10 | 7.41% |
| 25 | 10.82% |
| 50 | 14.19% |
| 100 | 18.56% |
| 250 | 26.44% |
| 500 | 34.75% |
| 1,000 | 44.95% |
| 2,500 | 61.05% |
| 5,000 | 74.60% |

The Gini coefficient of the distribution is 0.663: unequal, but not winner-take-all. A winner-take-all citation layer would put the first ten domains well past a quarter of the total. These ten hold less than a thirteenth.

The ten domains at the top, with their source-role classification from the release:

| Rank | Domain | Source role | Runs cited | Citation rate |
|---|---|---|---|---|
| 1 | reddit.com | Community and social platform | 2,012 | 12.75% |
| 2 | youtube.com | Search or media platform | 1,431 | 9.07% |
| 3 | linkedin.com | Community and social platform | 1,075 | 6.81% |
| 4 | medium.com | Editorial publication | 807 | 5.11% |
| 5 | forbes.com | Editorial publication | 649 | 4.11% |
| 6 | nih.gov | Academic and government source | 485 | 3.07% |
| 7 | gartner.com | Analyst and consulting research | 454 | 2.88% |
| 8 | techradar.com | Editorial publication | 448 | 2.84% |
| 9 | g2.com | Market database | 442 | 2.80% |
| 10 | arxiv.org | Academic and government source | 410 | 2.60% |

Six distinct source-role classes appear in those ten slots. There is no single class that owns the head.

## The median indexed domain is cited twice

The other half of a distribution is the part nobody publishes. Across 22,179 domains:

| Measure | Value |
|---|---|
| Maximum runs cited | 2,012 |
| 99th percentile | 49 |
| 90th percentile | 9 |
| 75th percentile | 4 |
| **Median** | **2** |
| 25th percentile | 1 |
| Mean | 5.0 |

- **10,381 domains — 46.8% of the index — were cited in exactly one answer run** across 125 days.
- 14,053 (63.4%) were cited in two runs or fewer.
- 18,166 (81.9%) were cited in five or fewer.
- Only 2,165 domains (9.8%) reach ten or more observed runs.
- Only 71 domains are cited in 100 runs or more.

The confidence grades follow directly. Of 22,179 domains, 14 carry grade A, 56 carry grade B, 430 carry grade C, and **21,679 are still collecting** — below the evidence floor, with no publishable rate. Being in the index is not an achievement; almost every domain the engines touched is in it. Clearing the floor is the achievement, and 500 domains have.

Breadth is just as thin. **17,186 domains — 77.5% — appear in exactly one segment** (one subject category paired with one question shape). The median domain appears in one. Only 710 domains appear in five or more. The maximum is 98.

Read together: the typical cited domain is cited once or twice, in one segment, and has not been observed often enough to earn a grade. It was a source for one answer to one kind of question, once.

## Concentration is a property of the source class, not of AI answers

This is the finding that changes what a publisher should do, and it is invisible in any single aggregate number. Split the 22,179 domains by their source role and measure concentration inside each class:

The release assigns every domain to one of nine source roles. Three of the nine hold fewer than thirty domains each, so the top-ten measure saturates inside them; the share held by each class's single most-cited domain is the measure that compares across all nine.

| Source role | Domains | Domain-run citations | Most-cited domain's share of its class | Top ten's share | Median domain |
|---|---|---|---|---|---|
| Search or media platform | 10 | 1,513 | **94.6%** (youtube.com) | whole class | 2.5 |
| Community and social platform | 27 | 4,314 | 46.6% (reddit.com) | 97.3% | 14 |
| Wire and press-release distribution | 9 | 226 | 42.9% (businesswire.com) | whole class | 5 |
| Analyst and consulting research | 364 | 3,381 | 13.4% (gartner.com) | 36.3% | 2 |
| Academic and government source | 413 | 3,872 | 12.5% (nih.gov) | 46.0% | 2 |
| Market and company database | 688 | 6,066 | 7.3% (g2.com) | 29.4% | 2 |
| Editorial publication | 1,222 | 13,309 | 6.1% (medium.com) | 25.7% | 2 |
| Vendor-owned source | 823 | 11,153 | **2.7%** (microsoft.com) | 18.8% | 3 |
| Other observed source | 18,623 | 67,043 | 0.4% (healthline.com) | 1.9% | 2 |

The nine rows account for all 22,179 domains and all 110,877 domain-run citations. Among the eight identified classes the top domain's share runs from 94.6% down to 2.7% — a thirty-five-fold difference in how much the leader of a class takes.

Search or media platforms look like a class and behave like one domain. Ten domains, 1,513 citations, and youtube.com holds 94.6% of them. The other nine were cited 82 times between them across 125 days; eight of those nine sit below the evidence floor and carry no grade at all. This is the most concentrated position measured anywhere in the release, and it is not a position a publisher can enter.

Community and social platforms are the next-tightest. The class has only 27 domains in it, and ten of those take 97.3% of everything the class is cited for. For a new community property, the class is effectively closed: the seats are taken by Reddit, LinkedIn, Substack, dev.to and Facebook, and the tail behind them is negligible.

The paid press-release layer is the smallest citation pool in the index. Nine wire and press-release domains were cited in 226 answer runs over 125 days — 0.20% of every citation recorded, less than the single tenth-ranked domain carries on its own (arxiv.org, 410). Three of the nine clear the evidence floor, all at grade C: businesswire.com in 97 runs, globenewswire.com in 71, financialcontent.com in 37. The other six have too little evidence to grade. Whatever the wire layer buys, appearing as a source in an AI answer is not measurably it.

Vendor-owned sources are the flattest of the identified classes, the residual bucket aside. 823 domains share 11,153 citations, the top domain takes 2.7%, and the median vendor domain is cited three times — the highest median of any class. A vendor page competes against a wide, shallow field rather than against an incumbent.

Editorial publications sit between: 1,222 domains, the largest classified citation pool at 13,309, and a leader on 6.1%. There is room in the class, but the median editorial domain is cited twice, which is the same as the median across the whole index. Being an editorial publication is not, by itself, an advantage.

The practical consequence: "how hard is it to get cited" has no single answer. It has nine answers, and which one applies to you is decided by what kind of source you are, not by how good the page is.

## The scarcest position is not volume, it is engines

There is a second concentration in this release, and it is steeper than anything in the volume curve. Split the 22,179 domains by how many of the six engines cited them at least once:

| Engines citing the domain | Domains | Share of all domain-run citations |
|---|---|---|
| 1 | 14,969 | 21.46% |
| 2 | 3,502 | 13.33% |
| 3 | 1,793 | 13.80% |
| 4 | 963 | 15.10% |
| 5 | 605 | 14.19% |
| **6** | **347** | **22.12%** |

347 domains are cited by all six engines. They carry more citation than the 14,969 cited by exactly one — 24,530 domain-run citations against 23,793. One domain cited by all six carries as much as forty-three cited by one. The median all-six domain was cited in 41 runs against a whole-index median of 2.

That group is not the media list the phrase suggests. Of the 347: 153 are unclassified, 85 vendor-owned, 58 editorial publications, 21 market and company databases, 13 analyst and consulting sources, 11 academic and government sources, 4 community platforms, 1 search or media platform, 1 wire service.

The engines differ enormously in how wide they cast. Over this window the release records citations to 9,497 distinct domains from Perplexity, 7,441 from Google AI Mode, 7,079 from Gemini, 5,138 from Claude, 4,877 from ChatGPT and 2,279 from Google AI Overviews. A domain cited by one engine has cleared one retrieval system's filter; a domain cited by six has cleared six filters that share very little machinery, each governed by its own user agent under RFC 9309.

This is where effort converts. Moving from one engine to two is a larger structural change than doubling volume on a single engine, and it is invisible in any citation-rate number.

## Inside a segment, the entry threshold is low

The whole-index curve overstates the difficulty, because no publisher competes against 22,179 domains at once. They compete inside a segment: one subject category paired with one buyer question shape. 85 of the 151 measurable segments in this release are published. Across them:

| Measure | Median published segment |
|---|---|
| Distinct domains cited in the segment | 231 |
| Share held by the segment's top ten | 24.1% |
| Share held by the segment's leader | 4.3% |
| Cited runs needed to reach the top 100 | 2 |
| Cited runs needed to reach the top 25 | 7 |
| Cited runs needed to reach the top 10 | 13 |
| Cited runs held by the segment leader | 34 |

Thirteen cited answers puts a domain in the top ten of the median segment, and seven puts it in the top twenty-five. Restricting to the 78 buyer-question segments and excluding the seven longer-running news segments, the number falls to twelve.

| Segment | Domains cited | Cited runs for top 10 | Leader's cited runs |
|---|---|---|---|
| AI Visibility & GEO, best tools | 351 | 17 | 30 |
| Cybersecurity, best tools | 213 | 13 | 44 |
| Enterprise Software, how buyers choose | 172 | 8 | 28 |

Segments vary far more than the median implies. The most open in the release is the legacy news topic, where 4,025 domains share citations and the top ten hold 8.3%. The most closed is Consumer Finance comparisons, where 170 domains share citations, the top ten hold 42.1%, and [acorns.com](https://www.acorns.com/) leads at a 38.41% citation rate. 20,704 of the 22,179 indexed domains — 93.35% — appear in at least one published segment, and 16,270 appear in exactly one of them. Most cited domains are cited about one thing.

## What this corrects

Machine Relations published an analysis on June 14, 2026 — [AI Citation Concentration: Why Market Databases Capture Disproportionate Share](https://machinerelations.ai/research/market-database-ai-citation-concentration-2026) — which described the distribution as a power law and read market databases as capturing disproportionate share. That analysis was run under MRI Score v1.1 over 7,124 domains and 28,870 source events, and it is labelled on the page as superseded by v2.0 on July 5, 2026.

At v2.0, over three times as many domains, the shape reads differently in two ways. The head is weaker than "power law" suggests in ordinary use: 1.81% for the leading domain, 7.41% for the top ten. And market databases are not the most concentrated class — their leader, g2.com, holds 7.3% of the class, which puts them sixth of the nine source roles, behind search or media platforms, community platforms, wire distribution, analyst research, and academic and government sources. The v1.1 reading was not wrong about market databases being strong in absolute terms; g2.com and crunchbase.com are rank 9 and rank 16 overall. It was wrong that their class is where concentration lives.

Correcting our own earlier published reading is part of what a versioned instrument is for. Every figure in this piece is tagged to release `mri_score_v2.0` generated 2026-09-18, and a later release can overturn it the same way.

## Where the independent research agrees

An empirical study of source coverage across LLM-based and traditional search engines, analysing 55,936 queries over six LLM search engines and two traditional engines, reports that LLM-based engines cite domains with *greater* diversity than traditional search engines, and that 37% of domains are unique to the LLM-based engines ([arXiv:2512.09483](https://arxiv.org/abs/2512.09483)). That is an independent measurement, on a different population with a different method, pointing the same way as the curve above: the citation layer of AI answers is wider than the intuition that it collapses to a few sites.

Two adjacent lines of work frame why the tail behaves as it does. A measurement framework distinguishing citation *selection* from citation *absorption* across AI search platforms argues that appearing as a citation and actually shaping the answer are separate dependent variables ([arXiv:2604.25707](https://arxiv.org/html/2604.25707v1)). And an observational audit of Google AI Overviews on "Your Money or Your Life" queries, using rank- and provenance-conditioned citation measures, finds citation behaviour that is not explained by retrieval rank alone ([Kakimov et al., PMLR v318](https://proceedings.mlr.press/v318/kakimov26a.html)). Neither study measures the concentration curve directly; both support the reading that source selection is its own process rather than a re-ranking of search results.

Access policy is the one structural factor a publisher directly controls. The Robots Exclusion Protocol is standardised as [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html); Google documents its crawlers and the `Google-Extended` token separately from Search inclusion ([common crawlers](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers), [crawler overview](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers)); OpenAI documents its own bots and their user agents ([OpenAI bots](https://platform.openai.com/docs/bots), catalogued independently at [Dark Visitors](https://darkvisitors.com/agents/gptbot)); Cloudflare has made per-site blocking of AI crawlers a one-click control ([Cloudflare](https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click/)). Google's own guidance on appearing in AI features is published in Search Central ([succeeding in AI search](https://developers.google.com/search/blog/2025/05/succeeding-in-ai-search)), and Search Console documents where AI-surface performance appears in its reports ([Search Console help](https://support.google.com/webmasters/answer/7576553)). Google has described AI Mode as a distinct search experience rather than a variant of the results page ([Google](https://blog.google/products/search/ai-mode-search/)). On the demand side, survey work finds users click source links less often when an AI summary is present ([Pew Research Center](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/)), which is why citation share and traffic share have to be measured separately. The broad web corpora these systems are built on are themselves public and enumerable ([Common Crawl](https://commoncrawl.org/)).

## What a publisher can check

Every one of the 22,179 domains in this release has a public page. A domain profile lists its overall citation rate, its confidence grade, its rank in the full universe, and every segment in which it has been observed — for example [reddit.com](https://machinerelations.ai/index/domains/reddit.com). Category and question-shape leaderboards list the top domains per segment, such as [AI infrastructure](https://machinerelations.ai/index/categories/ai-infrastructure). The [Index](https://machinerelations.ai/index) is the entry point, and the full release is machine-readable.

Domains that clear the evidence floor in a published segment also carry a citable badge stamped with this release id, served at `machinerelations.ai/index/domains/<domain>/badge.svg` and embeddable from the profile page.

Four checks are worth making against the numbers above. Which percentile of the curve your domain sits in, from the full-universe rank on its profile. Whether it clears the evidence floor, or is one of the 21,679 still collecting. Which source-role class it was assigned to, because the class decides which concentration figure applies to you. And how many of the six engines have cited it, because that number, not the rate, is what the steep part of this release's distribution is made of.

## Limits

This is a measurement of the Index's own observation window and question basket, not of the whole web. The Index runs a fixed basket of buying and research questions; a domain cited heavily for questions outside that basket will look thinner here than it is. Domains cited zero times in the window are not in the file at all, so the 22,179 is a population of *cited* domains, not of eligible ones — the true denominator of "sites that could have been cited" is larger and unknown.

Nothing here is causal. The curve describes what engines cited over 125 days. It does not say why, and it does not say that any intervention moves a domain along it. A domain's position at the next release is the only test of that.

Finally, concentration measured on domain-run citations is not the same as concentration measured on citation events, on answer position, or on whether the cited page actually supported the claim beside it. That last question is measured separately in [Answer-Source Fidelity](https://machinerelations.ai/measurement/answer-source-fidelity), and its results do not follow from this curve.

---

**Release:** `mri_score_v2.0`, contract `machine_relations_index_public_view_v2.0`, generated 2026-09-18. **Window:** 2026-05-10 to 2026-09-18, 125 days observed. **Population:** 22,179 cited domains, 15,782 observed answer runs, 110,877 domain-run citations. **Evidence floor:** 10 observed runs across 7 distinct run dates. See also the [Machine Relations Index Report for September 18, 2026](https://machinerelations.ai/research/machine-relations-index-report-2026-09-18) for the same release read by segment leader and engine.

## Attribution

This research is published by Machine Relations Research, the research program of machinerelations.ai — the public research and standards initiative that publishes the glossary, research, evidence, and measurements for the Machine Relations discipline. Provenance and editorial standards: https://machinerelations.ai/about

## Machine-readable related links

### Related concepts

- [Machine Relations Index (MRI)](https://machinerelations.ai/glossary/machine-relations-index)
- [MRI Score](https://machinerelations.ai/glossary/mri-score)
- [Machine Relations (MR)](https://machinerelations.ai/glossary/machine-relations)
- [AI Visibility](https://machinerelations.ai/glossary/ai-visibility)

### Supporting research

- [Ranked Without a Grade: Half of the AI Citation Index's Top-Ten Positions Belong to Ungraded Domains](https://machinerelations.ai/research/ranked-without-a-grade-ai-citation-index-domain-lookup-2026)
- [How to Rank in Perplexity: What the Citation Data Actually Shows](https://machinerelations.ai/research/how-to-rank-in-perplexity-citation-data-2026)
- [Entity Chain Scoring: How to Measure Cross-Domain Authority for AI Citation Eligibility](https://machinerelations.ai/research/entity-chain-scoring-measure-cross-domain-authority-2026)
- [How Stable Are AI Search Citations Week to Week? Evidence From Three Measurement Systems](https://machinerelations.ai/research/ai-citation-stability-week-to-week-evidence-2026)

### Framework context

- [Machine Relations Index](https://machinerelations.ai/index)
- [Machine Relations Stack](https://machinerelations.ai/stack)
- [Evidence Base](https://machinerelations.ai/evidence)
