Across the 16 categories in the Machine Relations Index that clear the evidence floor, the ten most-cited domains hold between 8.1% and 24.2% of citations, and reaching half of a category's citations takes between 49 and 247 different domains. There is no category in which a short list of 15 or 30 outlets covers most of what an answer engine cites. The concentration figure every pitch list is built on is a cross-topic aggregate, and at category level it does not survive.
What was measured #
Every figure here comes from the public release at machinerelations.ai/data/machine-relations-index.json, contract machine_relations_index_public_view_v2.0, release mri_score_v2.0+2026-09-23+fdec6388001a, methodology version mri_score_v2.0, generated 2026-09-23. The observation window runs 2026-05-10 to 2026-09-23 — 130 observed days, 16,475 observed answer runs across six engines (ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews and Perplexity), 23,280 cited source domains and 129,264 source events.
The Index scores each domain separately inside each subject category and each of six buyer-question shapes: best tools, how to choose, is it worth it, problem-first research, top lists, and head-to-head comparisons. For every domain and every category-and-shape cell it records runs_cited, the number of observed answer runs in which at least one engine cited that domain.
The unit of this study is the domain-shape-run citation: one domain, cited in one observed answer run, for one question shape inside one category. Summed over every domain's per-category cells, the 20 categories in the mri_taxonomy_v2.0 taxonomy hold 115,066 of them; the 16 categories this study reports hold 94,432. That unit answers the publisher's question — how often was this domain one of the sources for a question in my category — which is not the same as a raw citation-event count, because a single answer run can cite one domain more than once.
Every count below is read from the release, not re-derived by ranking. Per-stratum standing in the Index is published in relative_signal.category_signals[] and is not reconstructed by sorting, because the release uses competition ranking and a re-sort silently demotes tied domains.
The answer, by category #
Sorted by the number that matters to a pitch list: how many distinct domains you would have to be present on to cover half of the citations in your category.
| Category | Taxonomy id | Domains cited | Citations | Top 1 | Top 3 | Top 10 | Top 25 | Domains to half | Domains to 80% |
|---|---|---|---|---|---|---|---|---|---|
| Family Software | family-software |
643 | 3,424 | 3.83% | 9.40% | 24.21% | 39.49% | 49 | 221 |
| Consumer Finance | consumer-finance |
1,008 | 5,430 | 3.89% | 11.47% | 23.59% | 36.19% | 53 | 269 |
| AI Infrastructure | ai-infrastructure |
1,039 | 4,300 | 3.65% | 10.07% | 21.23% | 32.65% | 70 | 335 |
| Consumer Health | consumer-health |
1,304 | 6,062 | 3.74% | 9.70% | 18.84% | 29.53% | 95 | 417 |
| AI Security and Privacy | ai-security-privacy |
1,299 | 5,358 | 3.28% | 7.75% | 15.25% | 24.56% | 104 | 434 |
| Martech and Advertising | martech-advertising |
1,244 | 4,602 | 1.98% | 4.82% | 12.80% | 23.99% | 104 | 434 |
| Deep Tech and Hardware | deep-tech-hardware |
1,041 | 4,112 | 3.16% | 6.54% | 15.05% | 25.71% | 109 | 405 |
| Emergent Prosumer | emergent-prosumer |
867 | 3,229 | 4.00% | 8.05% | 14.96% | 23.69% | 118 | 375 |
| Education and Training | education-training |
1,076 | 3,878 | 3.58% | 6.45% | 14.23% | 23.41% | 123 | 437 |
| AI Visibility and GEO | ai-visibility-geo |
1,762 | 5,939 | 3.01% | 7.27% | 14.41% | 23.67% | 139 | 666 |
| Consumer Products | consumer-products |
1,544 | 5,757 | 3.14% | 7.73% | 13.65% | 21.35% | 157 | 579 |
| HR and Talent | hr-talent |
1,844 | 7,969 | 1.52% | 3.51% | 9.76% | 19.29% | 164 | 635 |
| Cybersecurity | cybersecurity |
2,161 | 9,573 | 1.91% | 5.46% | 12.16% | 20.03% | 175 | 699 |
| Fintech | fintech |
2,151 | 8,375 | 1.49% | 3.86% | 8.11% | 15.00% | 207 | 761 |
| Healthcare Services | healthcare-services |
2,078 | 8,218 | 2.51% | 5.27% | 10.56% | 17.47% | 214 | 786 |
| Enterprise Software | enterprise-software |
2,219 | 8,206 | 1.47% | 3.69% | 8.65% | 16.07% | 247 | 866 |
Four numbers in that table do work no aggregate figure can do.
No single domain reaches 4.1% of its own category. The most concentrated top position anywhere in the 20 categories is reddit.com at 4.00% of citations in Emergent Prosumer. In Enterprise Software the leader, linkedin.com, holds 1.47%. A category with a dominant source does not appear in this measurement.
The top 25 domains never reach 40%. The maximum across all 20 categories is 39.49%, in Family Software, the smallest and most concentrated category in the release. In Enterprise Software the top 25 hold 16.07%.
The spread between categories is five-fold. Family Software needs 49 domains to reach half; Enterprise Software needs 247. A quantity that varies five-fold across categories is not a constant, and a strategy calibrated to one category's curve is calibrated to the wrong number in every other.
Concentration is not a function of category size. Consumer Finance has 1,008 cited domains and reaches half in 53; AI Visibility and GEO has 1,762 and needs 139; but Consumer Health has 1,304 and needs 95 while Deep Tech and Hardware has 1,041 and needs 109. Breadth of the cited pool and the shape of its head move separately.
The claims this measurement tests #
Two concentration figures circulate widely enough to shape buying behaviour.
The first is that 15 websites own 68% of AI citations, which traces to the 5WPR AI Platform Citation Source Index, a synthesis of six earlier studies, announced via PR Newswire and reported by Everything PR. Our own provenance work on that figure, including its restatement on one of our pages, is at authoritytech.io.
The second is that roughly 30 domains own about two thirds of citations within any single topic, a framing repeated across vertical strategy guides such as AI+Automation's vertical playbooks and the vertical citation hub at ConnectEra, which reports single sources owning 85% or more of individual verticals.
Neither survives contact with a per-category measurement of this size. Among the 16 reported categories, the top 25 domains clear 30% in three and sit below 26% in 12. Reaching 80% of a category's citations takes more than 400 domains in 12 of the 16, and more than 700 in three.
This is not a claim that concentration is absent. It is a claim about where the concentration sits: in a thin head that no strategy can buy into, above a very long middle that is the only part a publisher can actually enter.
Where independent measurements land #
Two 2026 studies measuring the same quantity with different panels agree with this range, which matters more than either study alone.
Analyze's August 2026 study analysed 22,295 AI answers across ChatGPT, Perplexity and Google AI Mode, covering 115,843 citation events and 7,058 unique cited domains. It found the top ten domains covering 11% to 13% of citations per engine, and the top 50 covering 29% to 34%. GetIntel's August 2026 run logged 3,754 citations across ten software categories and five engines, going to 1,258 domains: the single most-cited domain took 3.5%, the top ten took under 15%, and 57% of domains were cited exactly once.
Both land inside our 8.1% to 24.2% band for the top ten, and GetIntel's 3.5% ceiling on a single domain matches ours at 4.00%. Three separate panels, three separate prompt sets, three separate engine mixes, one result.
The broader literature points the same way from other angles. Featured's June-August 2026 citation report found 34.5% of 22,881 citations going to domains with a Domain Authority under 40, across 11,499 different domains. Fahlout's synthesis of more than 30 studies records the four major engines sharing as little as 11% of their cited domains with one another, which is itself a reason a single cross-engine top list cannot describe any one engine's behaviour. Foglift's buyer-intent study found the cited set turning over almost completely as intent moves from discovery to shortlist, which is the same instability in a different axis.
Outside the commercial trackers, an academic result finds the same shape in a controlled setting. "When AI Writes, Who Gets Cited?", from teams at UT Austin, Stevens Institute of Technology, Washington University in St. Louis, Notre Dame and Rice, tested eleven models from three vendors on uniformly random panels and found the top decile receiving 23.3% to 30.2% of citations against 15.6% under indifferent selection. Sharp, measurable concentration — and still nothing like a handful of sources owning the answer.
Where sources disagree with this, the disagreement is usually a unit difference rather than a contradiction. Buffy's read of vertical concentration reports top-three brands holding more than 80% of AI visibility in some categories; that is share of answers mentioning a brand, not share of citations going to a domain, and the two can diverge completely because one brand's mention can be sourced from dozens of domains. Attrifast's 1,200-prompt vertical study, DeltaV's 25,000-citation study and Presenc's industry breakdown all report per-vertical differences in the same direction as ours, at smaller scale and with different category boundaries, which is why their absolute figures are not directly comparable to these.
What is excluded, and why #
Four of the release's 20 buckets are reported below the line rather than inside the finding.
Three of them — igaming-betting, industrial and local-services — carry no published stratum at all in this release. Every domain-and-question-shape cell inside them sits below the Index's evidence floor of 10 observations across 7 distinct run dates. Their curves are computed from observations that are individually too thin to publish a rate for, so their shape is dominated by the floor rather than by the market.
The fourth, legacy-unmapped, is not a subject category. It is the bucket for observations not mapped to a taxonomy node, and it is the largest bucket in the release at 17,037 citations across 4,025 domains. Its curve, with the top ten at 8.26% and 304 domains to reach half, is included here only because it is the widest and flattest distribution in the data and is the origin of the wider "49 to 304" range this measurement has been quoted at elsewhere on our properties.
| Excluded bucket | Domains | Citations | Top 10 | Domains to half | Reason |
|---|---|---|---|---|---|
legacy-unmapped |
4,025 | 17,037 | 8.26% | 304 | Not a subject category |
igaming-betting |
717 | 2,139 | 12.25% | 97 | No published stratum |
industrial |
704 | 1,057 | 7.00% | 188 | No published stratum |
local-services |
236 | 401 | 18.20% | 59 | No published stratum |
How robust the answer is #
The obvious objection to counting every observation is that the thin tail is noisy. So the same curves were recomputed using only cells the release marks published — those clearing the evidence floor on their own — and the answer barely moves.
In 13 of the 16 reported categories the published-only figures are identical, because every cell in them is published. In the three where they differ, Healthcare Services moves from a top ten of 10.56% to 10.94% and from 214 domains to half down to 199; HR and Talent moves from 9.76% to 9.53% and from 164 up to 170. The band across all 16 becomes 8.1% to 24.2% for the top ten either way.
That stability is the point. The finding is not an artefact of counting single observations, because the categories are wide enough that removing the unpublishable cells changes the answer by a few percent of itself.
Two limits belong on the record. First, these are domain-run citations inside a fixed taxonomy, and a category boundary is a modelling choice: a wider category will generally show a flatter curve, so the absolute counts belong to mri_taxonomy_v2.0 rather than to any universal definition of a market. Second, a 130-day window records what the engines did, not what they will do; the Index publishes drift series precisely because the composition of these heads moves between releases.
What a pitch list should do with this #
The practical consequence is not that concentration does not exist. It is that the reachable part of the distribution is much further down than the industry's working assumption, and it is category-specific.
Take the strongest possible version of a short list: the ten most-cited domains in a category, chosen with perfect foresight and including the platforms no comms team can pitch. In Enterprise Software that list covers 8.65% of the citations an engine has to choose from. In Family Software it covers 24.21%. Widen it to 25 domains and Enterprise Software reaches 16.07% while Family Software reaches 39.49%. Neither end of that is a strategy, and the distance between them is the whole reason a single cross-topic number was always going to mislead.
The number that does travel is the shape: a head of a few domains each holding under 4%, a middle of a few hundred domains that together hold the majority, and a tail of single observations. A publisher's realistic question is not which fifteen sources to join. It is which of the few hundred domains inside its own category's middle it can plausibly enter, and what the ordering of those looks like — which is what per-category, per-question-shape standing in the Index is for.
The whole-web version of this curve, across all 22,179 domains in the previous release, is at How Concentrated Are AI Citations?. Which kind of source wins inside each category, and how that changes with the buyer's question, is at Source Class Capture by Category.
Method notes #
- Counts are summed from each domain's
mri_score_v2.strata[]entries, grouped bycategoryacross all six question shapes. A domain contributes to a category once per shape in which it was cited. - "Domains to half" is the smallest number of domains whose citations, taken in descending order, reach 50% of that category's total. "Domains to 80%" is the same at 80%.
- The published-only sensitivity uses the same computation restricted to cells with
status: "published". - Shares are of the category's own total, not of the whole index, so they do not sum across categories.
- Categories and ids are the nodes of
mri_taxonomy_v2.0as carried in the release.