Fourteen public studies answer the question "how many sources does an AI answer cite" with figures from about 2 to nearly 22. That is not a narrow disagreement about engine behavior; it is evidence that "sources per answer" is being computed against at least two different denominators, and most of the spread disappears once the denominator is named. An audit of the Machine Relations Index's own September 2026 release found the identical trap sitting inside a public data field: a count named as if it were a per-answer citation tally is actually a distinct-domain count for an entire measurement stratum, and using it as a numerator understates true citation density by a factor of 3.2x.
The published range, and why it is not really about which engine cites more #
Read as a horse race between engines, the public figures look chaotic. One widely cited comparison puts Perplexity at 21.87 sources per answer against ChatGPT's 7.92 (Mentionova's citation-statistics roundup; the same 118,101-answer figure recurs at GEO Software Rankings). A different panel puts ChatGPT closer to Perplexity at 11.38 against 20.16 (Zerply's August 2026 monthly report). A third has Google AI Overviews citing more than either at 17.93, next to Gemini at 17.11 (Mentionova again, different metric on the same page). A fourth has AI Overviews at 3.2, Perplexity at 5.7 and ChatGPT at 2.1 — a completely different ordering and a completely different scale (FirstCitation's 527-response analysis). A fifth measures "distinct websites per answer" and finds all five major surfaces clustered tightly between 3.8 and 4.7 (Wellows).
None of these studies is simply wrong. Read closely, they are answering adjacent but different questions: total citation links versus distinct domains, per-query versus per-response, browsing-mode-only versus all responses including ones answered from the model's own knowledge with zero citations averaged into the denominator (on-page.ai's 793-citation study reports ChatGPT searching the web in only 21% of runs, which alone moves any all-run average far below any searched-run average). Query length changes the count on its own: one weglot.com breakdown shows Claude moving from 0.12 sources on a short query to 2.96 on a long one, and Google AI Overviews moving the opposite direction, from about 6 down to about 4 as queries lengthen (Weglot). PromptWatch's own longitudinal tracker shows a single engine's average swinging by nearly 10x within weeks as retrieval configurations changed under it (PromptWatch). SlateHQ separates "source links per answer" from "different sites cited," and the two numbers for the same engine differ by roughly 4x because one link total is spread across far fewer distinct domains (SlateHQ). LLM Pulse and Techshali each publish yet another pairing of numbers on a rolling window basis, none matching any of the above exactly (LLM Pulse; Techshali). Surfer's 2026 study of AI Overviews specifically reports 5 sources per query and, separately, that 99% of those sources are referenced only once per answer — a detail that matters below (Surfer SEO). DeepSmith and Lumen GEO both round up several of these figures into single scoreboards without reconciling the underlying denominators, which is the ordinary way this range gets restated as if it were one number (DeepSmith; Lumen GEO). AI Citation Monitor's Perplexity-versus-Google piece adds a sixth framing again, citations-per-answer against citation-inclusion-rate, two different things reported side by side as if comparable (AI Citation Monitor).
The unifying finding across all fourteen is not that any single number is correct. It is that a "sources per answer" figure is meaningless without naming what is being divided by what.
The same trap, found inside our own data #
The Machine Relations Index publishes a per-segment field, runs_cited_domains, alongside runs_observed, for every category-and-question-shape stratum in the public release (mri_score_v2.0+2026-09-25+04fcb7fb8f29, generated 2026-09-25, 131,670 source events across 23,872 cited domains). The field's name reads as a citation-event or run-level count. It is not. It is the number of distinct domains cited anywhere within that stratum across the whole 130-day observation window — a domain that appeared in one run counts exactly the same as a domain that appeared in fifty.
The same release carries the correct citation-event count in a different place: each domain's own record holds its per-stratum runs_cited value, the number of observed runs in which that specific domain was actually cited. Summing runs_cited across every domain's record for a stratum gives the true citation-event total for that stratum — a different, larger number than the top-level runs_cited_domains field for the same stratum, because most strata cite more than one domain per answer.
For the ai-infrastructure / best_x stratum specifically: the top-level field reads runs_cited_domains: 263 against runs_observed: 107. Summing runs_cited across every domain's record for that same stratum gives 872 citation events, not 263. Dividing the wrong field by runs observed gives 2.46 sources per answer for that one stratum; dividing the correct sum gives 8.15.
Portfolio-wide, across all 157 published strata: the top-level runs_cited_domains field sums to 36,248, against 16,788 total observed runs — a naive "sources per answer" of 2.16. The correct citation-event sum across every domain's per-stratum record is 117,271 — a true sources-per-answer figure of 6.99. The naive figure understates citation density by a factor of 3.2x, and the direction of the error is the same direction every one of the fourteen public studies above disagrees in: toward the smaller, more alarming-sounding number.
What an AI answer actually cites, by category #
Correcting the same error inside each of the 21 measured categories changes the ranking of nothing else in the release — the underlying runs_cited values were always right — but it changes the one number a buyer actually wants: how many source slots does a typical answer in my category carry.
| Category | Runs observed | Naive sources/answer (wrong field) | Correct sources/answer | Understatement |
|---|---|---|---|---|
| AI Visibility and GEO | 728 | 3.66 | 8.63 | 2.4x |
| Consumer Health | 786 | 2.08 | 7.71 | 3.7x |
| Martech and Advertising | 602 | 2.06 | 7.64 | 3.7x |
| AI Security and Privacy | 722 | 2.43 | 7.42 | 3.1x |
| Cybersecurity | 1,352 | 2.11 | 7.23 | 3.4x |
| Consumer Products | 798 | 2.30 | 7.21 | 3.1x |
| AI Infrastructure | 641 | 2.34 | 7.04 | 3.0x |
| HR and Talent | 1,164 | 2.06 | 6.89 | 3.3x |
| Consumer Finance | 798 | 1.65 | 6.80 | 4.1x |
| Fintech | 1,240 | 2.12 | 6.75 | 3.2x |
| Enterprise Software | 1,237 | 2.11 | 6.63 | 3.1x |
| Healthcare Services | 1,251 | 2.03 | 6.59 | 3.2x |
| Deep Tech and Hardware | 671 | 1.95 | 6.13 | 3.1x |
| Education and Training | 636 | 2.14 | 6.10 | 2.8x |
| Emergent Prosumer | 597 | 1.74 | 5.41 | 3.1x |
| Family Software | 642 | 1.36 | 5.33 | 3.9x |
Every category corrects upward, and the correction is never small: the smallest understatement in this table is 2.4x, in AI Visibility and GEO — the category this dataset itself sits in. A team benchmarking its citation coverage against the wrong field concludes it is competing for roughly two citation slots per answer; the release says the real number, in every one of these 16 categories, is between 5.3 and 8.6.
Why the error survives contact with the number #
A citation-count discrepancy this large would normally be caught by comparing two published figures, but nothing about either field is false. runs_cited_domains correctly counts distinct domains; it simply does not count what its name implies. This is the general failure mode behind the published range above: a study's headline number is usually arithmetically correct for the population it actually measured, and wrong only for the population its label promises. Surfer's finding that 99% of an AI Overview's sources are referenced only once inside a single answer is the fact that makes the two counts converge for that specific surface and diverge for others: the less repetition inside one answer, the closer a distinct-domain count and a citation-event count sit to each other, and the more a single answer repeats the same handful of domains across its component claims — common in longer, multi-part responses — the wider the two counts pull apart.
Anyone recomputing "sources per answer" from a public dataset that separates a distinct-count field from a per-record event count should sum the event-level field, never divide by the distinct-count field, and should say in the same sentence which of the two they used.
Methodology and sources #
Every figure attributed to Machine Relations in this piece is read directly from the public release at machinerelations.ai/data/machine-relations-index.json, contract version mri_score_v2.0, release mri_score_v2.0+2026-09-25+04fcb7fb8f29, generated 2026-09-25, covering a 130-day observation window across six engines (ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews and Perplexity), 23,872 cited domains and 131,670 recorded source events. The top-level strata array carries one record per category-and-question-shape combination with fields runs_observed and runs_cited_domains. Each domains[] record separately carries a mri_score_v2.strata[] array with a runs_cited field per category-and-question-shape combination for that specific domain. The portfolio and per-category figures above sum runs_observed and runs_cited_domains across the 157 top-level strata, and separately sum runs_cited across every domain record's per-stratum entries, matched by category and question shape. Both sums were computed directly against the release file rather than read from any intermediate report. The fourteen third-party figures cited above are quoted as published by their own sources, with the metric each one actually reports named alongside the number, per the methodology this page argues for.