AI answer engines systematically favor certain source types over others. Machine Relations Index v2 data across 17,533 domains and 11,009 observed answer runs shows that market databases place seven domains above a 1% citation rate, analyst research only two, and wire distribution services zero. The gap is not marginal — it is structural.
The Citation Hierarchy Across Seven Source Types #
The Machine Relations Index classifies every measured domain into one of seven source types based on its primary function: editorial publication, analyst research, market database, vendor-owned source, academic and government source, community and social platform, or wire distribution. Each type shows a distinct citation profile.
| Source Type | Domains Measured | Total Runs Cited | Avg Citation Rate | Confidence A Domains | Domains Above 1% |
|---|---|---|---|---|---|
| Community and social platform | 27 | 3,312 | 1.115% | 2 | 7 |
| Editorial publication | 1,088 | 11,050 | 0.093% | 2 | 13 |
| Vendor-owned source | 823 | 9,782 | 0.109% | 1 | 10 |
| Market and company database | 617 | 5,351 | 0.080% | 1 | 7 |
| Academic and government source | 360 | 3,152 | 0.081% | 2 | 4 |
| Analyst and consulting research | 364 | 3,034 | 0.077% | 1 | 2 |
| Wire and press-release distribution | 9 | 210 | 0.213% | 0 | 0 |
Community platforms — dominated by Reddit (12.90% citation rate) and LinkedIn (7.50%) — have the highest per-domain average because their small category has several extremely high-cited members. But the more operationally relevant comparison is among the source types that brands can directly influence: market databases, analyst research, vendor-owned content, and editorial placements.
This source-type hierarchy is consistent with independent research. A SurfacedBy analysis of 127,198 AI citations across five engines found that vendor, product, and long-tail sites accounted for 90.6% of citations on commercial-intent queries — with Reddit at just 1.8% and Wikipedia under 0.6%. A separate Binghamton University study of 366,087 citations documented that each AI engine develops distinct source-type preferences: OpenAI favors licensed publishers and wire services, Perplexity favors international editorial (BBC, The Guardian), and Gemini cites Forbes consistently across 11 B2B and B2C sectors.
Market Databases Lead in Competitive Depth #
Market and company databases — platforms like G2, Crunchbase, Grand View Research, and Fortune Business Insights — have seven domains cited in more than 1% of observed AI answer runs. That is more than analyst research (two domains), academic sources (four), and wire services (zero).
| Rank in Category | Domain | Citation Rate | Overall Rank | Confidence |
|---|---|---|---|---|
| 1 | g2.com | 3.22% | #8 of 17,533 | A |
| 2 | crunchbase.com | 2.46% | #15 | B |
| 3 | grandviewresearch.com | 1.71% | #22 | B |
| 4 | fortunebusinessinsights.com | 1.46% | #29 | B |
| 5 | mordorintelligence.com | 1.26% | #34 | B |
G2 — cited in 3.22% of all observed runs — ranks 8th among all 17,533 measured domains, ahead of every analyst research source except Gartner. The category has 617 total domains, but citation authority concentrates at the top: the five sources listed above account for 1,114 of the category's 5,351 total citations.
What makes this category distinct is its depth. Five market databases hold confidence tier B or higher. By comparison, analyst research has only two domains (Gartner and Deloitte) above the B threshold. Market databases offer AI engines something analyst firms often do not: structured, machine-readable comparison data across thousands of products and vendors, updated continuously.
Analyst Research Is the Most Top-Heavy Source Category #
Among 364 domains classified as analyst and consulting research, citation authority concentrates more sharply than in any other category. Gartner alone accounts for 389 of the category's 3,034 total citations — 12.8% of all analyst-source citations from a single domain.
| Rank in Category | Domain | Citation Rate | Overall Rank | Confidence |
|---|---|---|---|---|
| 1 | gartner.com | 3.53% | #6 of 17,533 | A |
| 2 | deloitte.com | 1.63% | #23 | B |
| 3 | mckinsey.com | 0.82% | #63 | C |
| 4 | pwc.com | 0.79% | #70 | C |
| 5 | galengrowth.com | 0.71% | #83 | C |
The dropoff from Gartner to McKinsey is a factor of 4.3x. McKinsey and PwC — firms with comparable or greater general brand authority — have citation rates below 1% and confidence tier C, meaning the evidence behind their rates is still relatively thin. Only 340 of the 364 analyst domains are still in the "collecting" stage, meaning 93% of classified analyst sources have not yet accumulated enough observations to receive a stable citation rate.
This top-heaviness has a practical implication. If a brand's strategy depends on getting cited alongside analyst research in AI answers, the available citation surface is extremely narrow: realistically Gartner and Deloitte, with a steep dropoff after.
Wire Distribution Faces a Structural Citation Ceiling #
Wire and press-release distribution is the only source type where no domain achieves a 1% citation rate. The entire category contains just nine classified domains, and its top performer — Business Wire at 0.77% — ranks 72nd overall. No wire distribution domain has earned confidence tier A or B.
| Rank in Category | Domain | Citation Rate | Overall Rank | Confidence |
|---|---|---|---|---|
| 1 | businesswire.com | 0.77% | #72 of 17,533 | C |
| 2 | globenewswire.com | 0.64% | #95 | C |
| 3 | financialcontent.com | 0.31% | #290 | C |
The 210 total citations across all wire services is less than what G2 alone accumulates (355 citations). This is consistent with earlier research on press release citation rates: syndicated press release content — identical text distributed across multiple wire endpoints — creates the kind of duplicate content that AI citation models are specifically designed to deduplicate. When every wire carries the same release, none becomes the canonical source an AI engine would prefer to cite.
The contrast with editorial publications is instructive. PR Newswire, which also publishes and distributes press content, is classified by the Machine Relations Index as an editorial publication rather than a wire service, and its 1.48% citation rate (ranking 27th overall) is roughly double that of Business Wire. The distinction appears to reflect PR Newswire's role as a content host — a destination where original announcements live — versus pure syndication infrastructure.
Editorial Publications Drive the Highest Total Citation Volume #
Editorial publications as a category generate the most total citations of any source type: 11,050 runs cited across 1,088 measured domains. The category spans everything from Medium (6.54% citation rate, ranked 4th overall) and Forbes (4.28%, ranked 5th) to hundreds of niche industry publications.
Only two editorial domains reach confidence tier A — Medium and Forbes. Twelve reach tier B, including Yahoo (2.64%), NerdWallet (2.13%), and TechRadar (2.02%). The remaining 1,007 domains are still in the collecting stage.
The 0.093% average citation rate across all editorial domains — lower than market databases, vendor-owned sources, or community platforms — reflects the category's long tail. Most editorial publications are rarely or never cited by AI engines. But the ones that are cited carry disproportionate weight: the top five editorial sources account for nearly 2,000 of the category's 11,050 total citations.
For brands, this means editorial placements are not interchangeable. A placement in a publication that AI engines already cite consistently (Medium, Forbes, or a top niche source in a relevant vertical) carries measurably different citation value than a placement in one of the thousand-plus editorial domains still in the collecting stage.
Vendor-Owned Sources Outperform Wire Services and Most Analyst Firms #
Vendor-owned sources — corporate blogs, documentation sites, and product pages — collectively generate 9,782 citations across 823 domains, making this the second-largest source type by total volume. Ten vendor-owned domains exceed a 1% citation rate, more than analyst research, academic sources, or wire distribution.
IBM leads the category at 2.74% (confidence tier A, ranked 10th overall). Microsoft follows at 2.51% (B, ranked 14th). Cybersecurity vendors like Palo Alto Networks (1.84%) and SentinelOne (1.63%) also rank in the top tier.
The structural takeaway: a vendor's own published content can earn higher AI citation rates than most analyst research and nearly all wire distribution. IBM's 2.74% citation rate is 3.3x higher than McKinsey's 0.82% and 3.6x higher than Business Wire's 0.77%. SurfacedBy's commercial-query analysis corroborates this: vendor, product, and documentation pages accounted for 90.6% of all citations in their commercial-intent sample, with Claude in particular drawing almost exclusively from documentation and vendor pages rather than editorial or community sources.
This challenges the traditional assumption that third-party validation inherently outperforms owned content in credibility signals — at least for AI answer engines, which appear to weight source structure, specificity, and update frequency alongside editorial independence.
Academic and Government Sources Anchor Factual Authority #
Academic and government sources have a distinctive citation profile: only 360 classified domains, but two of them — NIH (3.47%) and arXiv (3.16%) — reach confidence tier A. Both rank in the overall top 10.
Wikipedia (1.34%, tier B) and Harvard (1.21%, tier B) round out the tier, but the category drops steeply after that. Only 14 academic domains reach confidence tier C, and 342 are still collecting. The category's average citation rate (0.081%) is the second-lowest of any classified source type.
This pattern — extremely strong top anchors, then a sharp cliff — mirrors analyst research but with a wider top tier. For AI engines making factual claims about health, science, policy, or regulation, NIH and arXiv function as citation defaults. For other verticals, academic sources are cited opportunistically rather than systematically.
Why Source Type Matters More Than Individual Domain Authority #
The cross-type comparison reveals a pattern that individual domain profiles miss: AI engines are not simply ranking all sources on a single quality axis. They are selecting from source-type-specific pools based on the query's needs. A seven-month Conductor study of citation behavior across seven AI engines found that each engine develops a persistent "editorial identity" — a default source type it reaches for, intent by intent. Even Google's own products (AI Overviews, AI Mode, and Gemini) developed distinct source preferences for the same queries.
A query about enterprise security vendor comparison is structurally more likely to cite a market database (G2, Crunchbase) than an analyst firm, because market databases contain the structured comparison data the query demands. A query about macroeconomic trends affecting enterprise IT budgets is more likely to cite analyst research (Gartner, Deloitte), because those queries require synthesized industry analysis.
This means a brand's source architecture strategy should match the query types it wants to win. If the target queries are product comparisons and vendor evaluations, being cited or listed on G2 and Crunchbase matters more than a McKinsey mention. If the target queries are strategic and analytical, Gartner and Deloitte placements carry more weight.
Citation Concentration and the Long Tail Problem #
Across every source type, citation authority follows a power-law distribution. The top domain in each category captures a disproportionate share of that category's total citations:
| Source Type | #1 Domain | #1 Share of Category |
|---|---|---|
| Community and social | reddit.com (12.90%) | 42.9% of category citations |
| Analyst research | gartner.com (3.53%) | 12.8% |
| Academic and government | nih.gov (3.47%) | 12.1% |
| Market database | g2.com (3.22%) | 6.6% |
| Wire distribution | businesswire.com (0.77%) | 40.5% |
| Vendor-owned | ibm.com (2.74%) | 3.1% |
| Editorial publication | medium.com (6.54%) | 6.5% |
Two patterns emerge. In categories with few total domains (community, wire distribution), the top domain captures 40%+ of all citations — the category essentially IS that domain. In larger categories (editorial, vendor-owned, market database), the top domain's share is under 7%, meaning more domains compete for citation share.
For brands, this means the strategy differs by source type. Getting cited on Reddit is getting cited in the community category. Getting cited by a market database means competing against 617 other databases — but also means the opportunity is more broadly accessible, because multiple databases earn meaningful citation rates.
What This Means for Brand Source Architecture Strategy #
The MRI source-type data reshapes how brands should think about Machine Relations — the discipline of building and maintaining citation relationships with AI answer engines.
For companies selling enterprise software or services, the market database path is the most accessible high-citation channel. G2, Crunchbase, Grand View Research, Fortune Business Insights, and Mordor Intelligence all maintain citation rates above 1%. Getting listed, reviewed, or analyzed on these platforms creates citation surface across multiple AI engines simultaneously.
For companies seeking strategic positioning, analyst research carries concentrated authority — but the viable surface is narrow. Gartner and Deloitte are the only analyst domains with citation rates above 1%. A Gartner Magic Quadrant placement or a Deloitte industry report mention creates measurably higher citation probability than coverage from most other analyst firms.
For companies using press releases as a distribution strategy, the data suggests a structural problem. Wire distribution services collectively earn fewer AI citations than any single top-10 domain. A press release on Business Wire is cited in 0.77% of observed AI answer runs — a rate that makes it nearly invisible in most AI-generated answers. Brands relying on wire distribution for AI visibility may need to reconsider where their announcement content actually lives.
For companies with strong owned content, the vendor-owned category data is encouraging. IBM, Microsoft, Palo Alto Networks, and SentinelOne all earn higher citation rates than most analyst and academic sources. Original, structured content on a vendor's own domain can earn citation rates that compete with third-party editorial coverage.
Methodology: How the Machine Relations Index Measures Source Types #
The Machine Relations Index v2 measures citation rates across six AI answer engines — ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, and Perplexity — by counting how often each domain is cited in observed answer runs and dividing by total observed runs for the relevant segment. Citation rates are published only for domain-segment combinations that clear the evidence floor: at least 10 observations across at least 7 distinct run dates.
Each domain carries a confidence tier — A, B, C, or collecting — reflecting the volume of evidence behind its measured rate. Domains below the evidence floor are labeled "collecting" and do not receive a stable citation rate or ranking.
Source-type classification assigns each domain to one of seven categories based on its primary editorial function. The classification determines the comparison pool: a market database is ranked against other market databases, an analyst firm against other analyst firms. The overall ranking compares all 17,533 measured domains regardless of source type.
Data window for this analysis: May 10 through August 7, 2026 — 85 observed days across 11,009 answer runs.
FAQ #
Which source type gets cited most often by AI answer engines? #
Community and social platforms have the highest average per-domain citation rate (1.115%), driven by Reddit (12.90%) and LinkedIn (7.50%). Editorial publications generate the most total citations (11,050) due to having the largest domain pool (1,088 domains). Wire distribution has the lowest citation rates, with zero domains above 1%.
Why do wire services get cited less than other source types? #
Wire distribution services distribute identical content across multiple endpoints, creating duplicate content that AI citation models are designed to deduplicate. Business Wire's 0.77% citation rate — the highest among wire services — suggests that syndicated press releases rarely become the canonical source an AI engine selects for citation.
Are market databases or analyst firms better for AI citations? #
It depends on query type. Market databases have more domains above 1% citation rate (seven vs. two for analyst research) and broader competitive depth. Analyst research is more top-heavy — Gartner alone holds 12.8% of all analyst citations. For product comparison queries, market databases are more relevant; for strategic analysis queries, analyst firms carry more weight.
Can vendor-owned content compete with third-party sources for AI citations? #
Yes. IBM (2.74%), Microsoft (2.51%), and Palo Alto Networks (1.84%) all earn higher citation rates than McKinsey (0.82%), PwC (0.79%), and nearly all wire services. Original, structured vendor content can earn citation rates that exceed most third-party analyst and editorial sources.
How does the Machine Relations Index classify source types? #
Each of the 17,533 measured domains is assigned to one of seven categories — editorial publication, analyst research, market database, vendor-owned source, academic/government source, community/social platform, or wire distribution — based on its primary function. The classification determines intra-type rankings while the overall ranking compares all domains together.
What confidence tier do I need to trust a citation rate? #
Confidence tier A domains have the strongest evidence base — enough observations to produce a reliable rate. Tier B has moderate evidence. Tier C has limited evidence. Domains marked "collecting" have not yet cleared the evidence floor of 10 observations across 7 distinct run dates and do not have a settled rate.