Last updated: July 27, 2026
ChatGPT cites 3,015 distinct domains across the queries the Machine Relations Index measures — roughly 18% of the 16,356 domains observed across all six engines. Getting into that set is not random. The data shows specific source types, content structures, and domain characteristics that predict whether ChatGPT selects a source for citation.
This analysis uses MRI v2 citation rate data from a 74-day measurement window (May 10 – July 26, 2026) across 9,498 observed runs and six AI answer engines, combined with external research on ChatGPT's retrieval pipeline.
How ChatGPT Retrieves and Cites Sources #
ChatGPT does not cite from its training data. When a user asks a question that triggers web search, the model generates multiple search queries from a single prompt, retrieves 20–40 URLs from Bing's index, extracts relevant passages, then synthesizes an answer with inline citations pointing to the sources those passages came from.
The citation goes to the passage that made it into the synthesis, not to the page with the highest search ranking. This is a meaningful distinction: a page can rank first in Bing but never get cited if its content is not structured for passage extraction. Kime.ai's analysis found that ChatGPT cites only about 15% of the pages it retrieves, and 44.2% of citations come from the first 30% of a page — making opening content disproportionately important.
QuickSEO's cross-engine study measured ChatGPT at an average of 6.88 sources per response (versus Perplexity's 16.35), though Semrush's 2026 AI Visibility Index reported a higher average near 15 sources across 126 million prompts. The difference likely reflects query complexity: buying-decision queries pull more sources than simple factual lookups. Across 34,000 monitored answers, ChatGPT cited sources in 87% of responses.
Which Source Types ChatGPT Cites #
The Machine Relations Index classifies every cited domain into one of nine source roles. The citation distribution reveals where ChatGPT — alongside five other engines — draws its references.
| Source Role | Domains | Citation Volume | Citations per Domain |
|---|---|---|---|
| Editorial media | 1,035 | 9,884 | 9.5 |
| Vendor-owned | 823 | 9,345 | 11.4 |
| Market database | 596 | 5,165 | 8.7 |
| Analyst research | 364 | 3,000 | 8.2 |
| Community/social | 27 | 2,818 | 104.4 |
| Academic/government | 356 | 2,817 | 7.9 |
| Platform search | 10 | 932 | 93.2 |
| Wire distribution | 9 | 203 | 22.6 |
Source: Machine Relations Index v2, 16,356 domains measured, 74-day window. Citation volume = total runs where at least one answer cited a domain in that role.
Two patterns stand out. Editorial publications and vendor-owned websites each account for roughly 28–29% of categorized citation volume — vendor-owned domains nearly match editorial media. And community platforms (primarily Reddit, LinkedIn, Substack, and dev.to) have extreme citation density: 27 domains generate as much citation volume as 356 academic sources.
LLM Pulse's analysis of 9.6 million queries confirmed this concentration: the top 30 domains capture roughly 67% of citations within a topic, and the top 10 account for 46%. But the long tail is real — 58% of cited URLs appear only once, meaning smaller domains can earn citations on specific queries even against dominant sources.
What the Top-Cited Domains Have in Common #
The 30 highest-cited domains in the MRI span every source role. Here are the top 15:
| Rank | Domain | Source Role | Citation Rate | Engines |
|---|---|---|---|---|
| 1 | reddit.com | Community | 11.81% | 4 |
| 2 | youtube.com | Platform | 9.00% | 6 |
| 3 | linkedin.com | Community | 7.99% | 6 |
| 4 | medium.com | Editorial | 6.97% | 6 |
| 5 | gartner.com | Analyst | 4.00% | 6 |
| 6 | forbes.com | Editorial | 3.92% | 6 |
| 7 | g2.com | Market database | 3.62% | 6 |
| 8 | arxiv.org | Academic | 3.30% | 6 |
| 9 | ibm.com | Vendor-owned | 3.01% | 6 |
| 10 | nih.gov | Academic | 3.07% | 6 |
| 11 | crunchbase.com | Market database | 2.85% | 6 |
| 12 | microsoft.com | Vendor-owned | 2.47% | 6 |
| 13 | yahoo.com | Editorial | 2.41% | 5 |
| 14 | techradar.com | Editorial | 2.18% | 3 |
| 15 | substack.com | Community | 2.14% | 5 |
Source: Machine Relations Index v2. Citation rate = runs where the domain was cited / total observed runs. Engines = number of the six measured engines that cited this domain at least once.
The top 15 include five editorial publications, four community platforms, two vendor-owned websites, two market databases, one analyst firm, and one academic repository. Cross-engine citation coverage correlates with stability: domains cited by all six engines tend to maintain citation rates across measurement periods, while domains cited by fewer engines show more volatility. Citation patterns can shift abruptly — LLM Pulse documented Reddit citations collapsing from roughly 60% to 10% of responses in September 2025 after Google removed a SERP parameter, demonstrating how fragile single-engine citation authority can be.
Only 264 of 16,356 measured domains (1.6%) are cited by all six engines. Of those, 74 are vendor-owned — meaning a brand's own website, if it reaches that threshold, has demonstrated citation authority across the entire engine landscape.
Five Factors That Drive ChatGPT Citation #
Combining MRI measurement data with external research on ChatGPT's retrieval behavior, five factors consistently predict citation.
1. Answer-First Content Structure #
A study of 1,000 websites and 180,000+ citations found that pages placing direct answers in the first 100–150 words received 67% more ChatGPT citations than pages that buried the answer. ChatGPT extracts passages, not pages. If the answer is below the fold or wrapped in preamble, the passage extractor skips it.
2. Original Research and Data #
The same study found original research and data studies received 3.2× more citations than aggregated content. The GEO-bench study (KDD 2024) confirmed this: adding source citations and statistics to content increased AI visibility by up to 40%. LLM Pulse's analysis found that 52.2% of cited blog posts featured original surveys, benchmarks, or datasets.
This aligns with MRI data: market databases like G2 (citation rate 3.62%), Crunchbase (2.85%), and Grand View Research (1.98%) all rank in the top 30. These are structured data sources, not editorial narratives.
3. Bing Indexation and Crawl Access #
ChatGPT retrieves from Bing's index, not Google's. A page that ranks in Google but is not indexed in Bing will never appear in a ChatGPT response. Blocking GPTBot or OAI-SearchBot in robots.txt eliminates a domain from consideration entirely.
DigitalApplied's 2026 data study rated URL accessibility as the highest-evidence citation factor at 9.5/10. Preview control (nosnippet directives) scored 9.2/10 for its ability to suppress citations — confirming that crawl policy directly controls citation eligibility.
4. Topical Focus Over Breadth #
Sprinklr's SIGIR 2026 research found that tightly focused pages — the kind that answer one question well — were cited far more often than broad "ultimate guide" content. A targeted 600-word answer outperformed comprehensive 4,000-word pillar pages. Seven of eighteen tested content optimizations showed weak or inconsistent impact, meaning most formatting and structural tricks do not move the needle.
5. Cross-Source Corroboration #
ChatGPT preferentially cites claims that appear across multiple sources. The cite.solutions analysis describes a "trust gate" where claims corroborated across independent sources win citation slots over marketing-only claims from a single vendor. SolCrys's analysis of 17,551 citations found that corroboration across 2–3 authoritative sources is a consistent citation predictor, and that the first-citation lag for newly eligible pages is typically 24–72 hours but actual citation events usually take 4–12 weeks as corroboration builds. DigitalApplied's data found branded web mentions correlate 3× more strongly with AI visibility than backlinks (0.664 vs. 0.218 correlation).
What Does Not Improve ChatGPT Citation #
Not every popular recommendation holds up under measurement.
Formatting tricks have negligible effect. Sprinklr's SIGIR research tested reordering headers, restructuring bullets, and adding "AI-friendly" formatting. None had measurable impact on citation rates.
The "4.3× freshness premium" lacks a primary source. This widely cited claim does not trace to a published study. DigitalApplied's analysis found the actual freshness advantage is more modest: AI-cited content averages 1,064 days old versus 1,432 days for organic results — a 25.7% difference, not a multiplier. Freshness matters, but it is not the dominant signal many guides claim.
Wire distribution is nearly invisible. MRI data shows wire and press-release distribution services account for 0.6% of categorized citation volume across nine domains. DigitalApplied measured press releases at 0.04% of AI citations. Distributing a press release does not earn ChatGPT citations.
ChatGPT and other engines cite different sources. QuickSEO's cross-engine analysis found that only 12% of URLs cited by AI tools overlap with Google's top 10 results, and the top 15 domains capture 68% of all consolidated AI citation share. ChatGPT leans heavily on Wikipedia (26–48% of top-10 citations) and Bing-indexed sources, while ScaleGrowth's research found that only 30% of brands remain visible in back-to-back AI responses for the same query — demonstrating how volatile citation positions are even for established domains.
Citation as a Measurable Channel #
ChatGPT citation is not a side effect of SEO. It is a distinct discovery channel with its own retrieval pipeline, its own index dependency (Bing), and its own source selection criteria. The Machine Relations framework treats each AI engine's citation behavior as a measurable surface — tracked through citation rates, source roles, and engine-specific source selection differences.
The MRI measures ChatGPT alongside Perplexity, Gemini, Claude, Google AI Mode, and Google AI Overviews. Domains that earn citations across multiple engines demonstrate broader source authority than those cited by a single engine. Only 1.6% of measured domains cross the all-engine threshold — and the characteristics that drive multi-engine citation (original data, answer-first structure, corroboration from independent sources) are the same ones that predict ChatGPT citation specifically.
For brands evaluating their AI visibility strategy: ChatGPT citation is not won through tricks or formatting. It is won through the same fundamentals that build source authority everywhere else — useful content, real data, proper indexation, and independent corroboration.
FAQ #
Does ChatGPT use Google's index for citations? #
No. ChatGPT retrieves from Bing's web index. A page indexed in Google but missing from Bing will not appear in ChatGPT responses. Ensuring Bing indexation and allowing GPTBot/OAI-SearchBot in robots.txt are prerequisites for citation eligibility.
How many sources does ChatGPT cite per response? #
ChatGPT averages approximately 15 sources per response according to Semrush's 2026 AI Visibility Index. Not every response includes citations — the 87% citation rate means roughly 13% of responses use training knowledge without web retrieval.
Can a vendor's own website get cited by ChatGPT? #
Yes. Vendor-owned websites account for 27.4% of categorized citation volume in the Machine Relations Index — nearly matching editorial publications at 28.9%. Seventy-four vendor-owned domains are cited by all six measured engines, including ChatGPT. A brand's own website is one of the largest citation surfaces available.
Do press releases help with ChatGPT citations? #
Press release wire services account for 0.6% of categorized citation volume in MRI data. Independent measurement found press releases at 0.04% of AI citations. Wire distribution does not meaningfully contribute to ChatGPT citation.