Research

The Largest Source Class in the Head of the Index Is Companies' Own Websites

In the September 25, 2026 release of the Machine Relations Index, 29 of the 100 most-cited domains are vendor-owned sources, against 22 editorial publications, 16 market and company databases and one wire domain. Across the whole index a vendor-owned domain averages 13.9 domain-run citations; an editorial publication averages 10.8.

Published Machine Relations Research
Index Analysis
TopicsAI SearchCitationsMeasurementSource SelectionOwned Media

The Machine Relations Index publishes a standing table of the hundred most-cited source domains on the open web. In the September 25, 2026 release, twenty-nine of those hundred domains are companies' own websites. Twenty-two are editorial publications. Sixteen are market and company databases, seven are community and social platforms, seven are academic or government sources, six are analyst and consulting research, and one is a wire and press-release distribution domain.

A company's own site is the most common kind of source in the head of the index. It is more common there than the press.

That is a measurement of one hundred domains, and it is worth stating what it is not: it is not a share of citations, and the head of this index is a thin slice of it. Both of those are addressed below. But the count is the finding, because the budget question it bears on — whether AI visibility is bought through placement or built on a domain you control — is usually answered from the opposite assumption.

The head of the index, by source class #

Source class Domains in the top 100
Vendor-owned source 29
Editorial publication 22
Market and company database 16
Other observed source 11
Community and social platform 7
Academic and government source 7
Analyst and consulting research 6
Search or media platform 1
Wire and press-release distribution 1

The three most-cited domains overall are reddit.com, youtube.com and linkedin.com. The first vendor-owned domain appears at standing #12.

The twenty-nine, in standing order, are microsoft.com, ibm.com, landbase.com, paloaltonetworks.com, sentinelone.com, guideflow.com, stripe.com, hubspot.com, salesforce.com, experian.com, rippling.com, nvidia.com, zapier.com, databricks.com, monday.com, huntress.com, crowdstrike.com, oracle.com, truefoundry.com, semrush.com, wiz.io, atlan.com, improvado.io, pr.co, airwallex.com, redhat.com, agilitypr.com, cynet.com and 6sense.com.

Nineteen of the twenty-nine carry a confidence grade of A or B in this release; the remaining ten carry C, the lowest published grade, meaning the rate stands on less evidence than the rows above it. The list is a mix of the largest software companies in the world and companies most readers will not recognise. That specific point — that citation standing in this class does not track company size — was measured on this index in June, on the superseded v1.1 methodology, and is reported there. It is not the finding here and it is not re-argued.

Across the whole index, not just the head #

The head is one hundred rows out of 23,872 cited domains. Across all of them, the release classifies every cited domain into exactly one of nine source roles. The nine close on the full universe:

Source class Domains Domain-run citations Share Mean per domain
Other observed source 20,225 72,101 61.5% 3.6
Editorial publication 1,288 13,926 11.9% 10.8
Vendor-owned source 823 11,403 9.7% 13.9
Market and company database 700 6,244 5.3% 8.9
Community and social platform 27 4,435 3.8% 164.3
Academic and government source 426 3,967 3.4% 9.3
Analyst and consulting research 364 3,439 2.9% 9.4
Search or media platform 9 1,509 1.3% 167.7
Wire and press-release distribution 10 247 0.2% 24.7

Editorial publications take more total citations than vendor-owned sources — 13,926 against 11,403 — but they do it with 57 percent more domains. Per domain, the vendor-owned class is cited more: a mean of 13.9 domain-run citations against 10.8. Among the seven classified classes that hold more than ten domains, vendor-owned has the highest mean of any class except the two platform classes, which are 27 and 9 domains wide and are structurally different animals.

The two platform classes are the real concentration story in this table, and it is not a story a brand can act on. Twenty-seven community and social platforms and nine search or media platforms hold 5.1 percent of all domain-run citations between them, at means of 164.3 and 167.7. Nobody buys their way into that class.

The wire and press-release distribution class holds ten domains and 0.2 percent. That result is consistent across every cut of this index it has been measured on, including the news-driven segments built specifically for breaking-news queries, where across seventy top-ten slots the wire class holds none.

Where this sits against the published consensus #

The widely repeated figure in this market is that earned media accounts for roughly 84 percent of AI citations and brand-owned content for roughly 16, a reading of Muck Rack's May 2026 Generative Pulse study that has been restated as a rule of thumb across the AI-visibility category, and which our own July synthesis of nine independent studies carried forward. Practitioner write-ups have since noted how far the estimates diverge — one review of eight studies puts the earned-media share anywhere between 40 and 85.5 percent depending on who counts and what counts as earned.

Our measurement does not refute that range and is not the same measurement. A share of citation events and a count of domains in a head are different quantities, and studies differ on whether platforms like Reddit and YouTube are counted as earned media at all — a choice that moves the headline number more than any other single decision. Other measurements of the same question land in the same territory of disagreement: Profound and Geonimo publish their own source-mix breakdowns, one analysis of 22,295 answers found that the top ten domains cover only 12 percent of B2B citations, and Semantica and Ranqo publish different most-cited rosters again. Academic work on how models source brand information across markets is early.

What the count above does say is narrower and harder to wave off: if you open the standing table of the hundred domains this index observes most, the largest group in it is companies writing on their own domains, and the class that exists to distribute announcements is one row.

What this changes for a budget #

A PR strategy aimed only at editorial placement is aimed at 22 of the 100 rows in the head of this index, and at a class whose per-domain citation rate is below that of the companies it writes about. That is an argument for adding an owned-source line to the plan, not for deleting the earned one — editorial publications are still the second-largest group in the head and still take more total citations than any other classified class. The two do different work.

It is also not an argument that publishing more pages produces citations. The vendor-owned class holds 823 domains. Twenty-nine of them are in the head. The other 794 are not, and this measurement says nothing about what separates them beyond what is visible in the table: the ones in the head are cited across the engine set on the buying questions their categories are measured on.

What this measurement does not say #

These counts are floors, not ceilings. Two Google surfaces had collection gaps inside this window: Google AI Overviews recorded no citations from 2026-08-15 to 2026-09-24, and Google AI Mode recorded none from 2026-09-14 to 2026-09-24. ChatGPT, Claude, Gemini and Perplexity collected throughout. Every citation count above is therefore at least the number shown. Unlike a comparison across segments measured on one identical window, a comparison across source classes is only unbiased if those two surfaces cite the nine classes in roughly the proportions the other four do. Nothing in this release establishes that, so the class composition above should be read as the composition the collected set produced, not as a constant.

Source-role labels are downstream annotations, not ranking inputs. They cover part of the index: 20,225 of 23,872 cited domains fall into "other observed source", which is the unclassified remainder and holds 61.5 percent of domain-run citations. Every class comparison here is a comparison among the 3,647 classified domains. A domain moving between classes would move these counts, and the sensitivity of that classification has been measured separately.

"Domain-run citations" is the quantity the source-role table reports: one domain cited in one observed answer run, counted once. It is not the release's citation-event count of 131,670, which counts each cited URL. The nine classes sum to 117,271 domain-run citations, and every percentage in the class table is taken against that total rather than against the event count.

This index measures which sources answer engines cite when the market's buying questions are asked. It does not measure whether a placement caused a citation, and no figure here should be read that way.

Methods #

Every figure comes from the public release at machinerelations.ai/data/machine-relations-index.json, contract machine_relations_index_public_view_v2.0, methodology mri_score_v2.0, release mri_score_v2.0+2026-09-25+04fcb7fb8f29, artifact 04fcb7fb8f29, generated September 25, 2026. The window runs May 10 to September 25, 2026, 132 days observed: 16,788 answer runs, 1,010 monitored prompts, 131,670 citation events, 23,872 cited source domains, six answer engines.

The six are ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews and Perplexity. Each documents its own retrieval and citation surface: OpenAI documents a search crawler distinct from its training crawler at OpenAI bots, Anthropic documents citations as part of Claude's web search, Google documents grounding with Search for Gemini and AI features and your site for its search surfaces, and Perplexity documents its own retrieval surface in its developer overview. Each reaches the web under its own user agent, governed by the Robots Exclusion Protocol as standardised in RFC 9309.

The head-of-index table is the Index's published standing table of the hundred most-cited domains, read at machinerelations.ai/index. That table labels editorial publications by their standing within the editorial class rather than their overall standing; the hundred rows are nonetheless the overall top hundred, and the twenty-two class-labelled rows occupy the twenty-two overall positions the other seventy-eight rows leave open. Class counts above are counts of those hundred rows.

Citation rates are published only for segments clearing the Index's evidence floor of at least 10 observed runs across at least 7 distinct run dates. Domain confidence grades A, B, C and collecting reflect the volume of evidence behind a domain's rate.

Machine Relations publishes this index as neutral public research. The full methodology, the crosswalk from the retired v1.1 scoring and every published segment are at machinerelations.ai/index.