# Why AI Citation Studies Disagree: The 2026 Source Register

A source-by-source audit of the AI citation statistics most often repeated in GEO and PR, showing why 25%, 39.5%, 47%, 65%, 84%, and 85.5% are not competing estimates of one thing.

Canonical URL: https://machinerelations.ai/research/why-ai-citation-studies-disagree-source-register-2026
Published: 2026-09-09
Research type: Research Synthesis

## Source Body

# Why AI Citation Studies Disagree: The 2026 Source Register

AI citation studies do not agree because they are usually not measuring the same population, source taxonomy, engine set, query set, or denominator. The widely repeated figures—25%, 39.5%, 47%, 65%, 84%, and 85.5%—cannot be ranked as if they were estimates of one universal “earned media share.” One is journalism inside a 25-million-link vendor study. One combines earned and news across eight engines. One is a single-topic, single-platform study. One is a provider-specific maximum. One uses an expansive third-party-source taxonomy. One has no matching primary Muck Rack result.

The Machine Relations rule is simple: **do not compare a number until you can name its measured unit**.

## The source register at a glance

| Published figure | Source and study | What was measured | What the figure can support |
|---|---|---|---|
| **25–27% journalism** | [Muck Rack, *What Is AI Reading?*](https://muckrack.com/blog/what-is-ai-reading-may-2026) | Cited links in responses from ChatGPT, Claude, and Gemini; May 2026 edition analyzed more than 25 million links across 17 industries | Journalism is a material source class in that study population. It does not measure all AI visibility or the causal effect of media relations. |
| **82–89% broad non-owned, non-paid sources** | [Muck Rack’s three-edition explanation](https://www.prdaily.com/what-does-earned-mean-in-the-age-of-ai/) | Journalism plus encyclopedic, government, academic, social, corporate/blog, and other third-party sources; Muck Rack labels the broad pool “earned media” | AI answers in the measured engines rely heavily on sources brands do not own or pay for. It cannot be glossed as journalism or media placements alone. |
| **39.5% earned/news** | [Meltwater, March–April 2026 AI visibility research](https://www.meltwater.com/en/blog/ai-search-visibility-march-april-2026) | Citations observed by Meltwater across a different engine and category mix; the April cut reports Earned/News as 39.5% | Earned and news sources are substantial in Meltwater’s taxonomy and panel. The result is not a direct contradiction of Muck Rack because the categories and engines differ. |
| **47% third-party news and informational sources** | [Fullintel–UConn study preview](https://fullintel.com/blog/ai-search-citations-news-websites/) | 400 prompts, 10 personas, weight-loss-drug queries, citations surfaced through Scrunch AI | News and credible informational sources drove nearly half the citations in this narrow health-communications study. It is not a multi-engine universal benchmark. |
| **65% earned-media share for Claude; 46% for Gemini** | [Chen, Wang, Chen, and Koudas, *Navigating the Shift*](https://arxiv.org/abs/2601.16858) | Comparative web-search and generative-answer source composition, reported by provider | Source composition varies sharply by provider. A provider maximum should not be restated as an all-engine rate. |
| **84% “earned media”** | [Muck Rack, May 2026 edition](https://muckrack.com/blog/what-is-ai-reading-may-2026) | Muck Rack’s broad third-party category, with journalism separately reported at 27% | The measured answers relied heavily on non-owned, non-paid source classes. Muck Rack itself says this does not mean media relations drives 84% of AI visibility. |
| **85.5%** | [5W’s *AI and the Israeli Brand*](https://www.5wpr.com/research/ai-israeli-brand/) and [5W-issued release](https://www.prnewswire.com/news-releases/85-5-of-ai-citations-come-from-earned-media--not-brand-websites-5w-releases-ai-and-the-israeli-brand-mapping-the-new-discovery-funnel-302771336.html) | 5W publishes the figure while attributing it onward to Muck Rack | It can be described as a figure published by 5W. It should not be described as a verified Muck Rack finding: the reviewed Muck Rack editions report 89%, 82%, and 84%, not 85.5%. |

These numbers are evidence of **measurement diversity**, not statistical consensus.

## The 84% claim is a taxonomy problem before it is a strategy claim

Muck Rack’s May 2026 edition is large: more than 25 million cited links from ChatGPT, Claude, and Gemini responses across 17 industries. It reports 84% under a broad “earned media” category and 27% as journalism.

The gap matters. Muck Rack’s broad category includes journalism, but also encyclopedic sources, government sites, academic sources, social platforms, corporate blogs, and other third-party material. In a September 2026 clarification, Muck Rack’s data and communications leaders wrote that they would not tell a communications leader that media relations drives 84% of AI visibility. Their study does not say that.

The evidence-safe interpretation is narrower and more useful: in Muck Rack’s measured population, sources outside the focal brand’s owned and paid channels supplied most cited links. Journalism supplied roughly one quarter. That supports a source-portfolio strategy. It does not support attributing the entire 84% to press coverage, placements, or PR causality.

## The 85.5% figure fails the source-existence check

The 85.5% figure appears on 5W research pages and in a 5W-issued PR Newswire release. Those pages attribute the number to Muck Rack.

The reviewed Muck Rack editions do not publish 85.5%. The July 2025 edition reports 89% under the broad earned category, the December 2025 edition reports 82%, and the May 2026 edition reports 84%. Journalism remains a separately reported 25–27% across the editions.

That produces a precise attribution boundary:

- **Verified publisher of the exact 85.5% figure:** 5W.
- **Onward attribution made by 5W:** Muck Rack.
- **Verification status of that onward attribution:** unsupported by the reviewed Muck Rack editions.

This does not prove how 5W calculated or inherited the number. It proves only that a live page can carry an exact statistic and a named source while the named source does not contain that statistic.

## Why 39.5%, 47%, and 65% are not rebuttals to 84%

Meltwater reports 39.5% for an Earned/News category in its April 2026 analysis. Fullintel and UConn report 47% third-party news and credible informational sources in a study limited to 400 prompts, 10 personas, weight-loss-drug questions, and one commercial AI-search platform. Chen and colleagues report provider-specific source composition, including a 65% earned share for Claude and 46% for Gemini.

Those findings should not be averaged. They should not be used to vote on the “real” percentage. Each belongs to a different measurement contract:

1. **Different engine sets.** A panel containing ChatGPT, Claude, and Gemini is not the same instrument as an eight-engine panel, a one-platform interface, or a provider-by-provider comparison.
2. **Different prompts and domains.** Cross-industry brand prompts, B2B categories, health queries, and general informational queries produce different source mixes.
3. **Different source taxonomies.** “Journalism,” “news,” “earned,” “third party,” “informational,” and “non-owned/non-paid” are not interchangeable labels.
4. **Different denominators.** A cited link, an answer containing a citation, a prompt, a source domain, and a brand mention are different units.
5. **Different time windows.** Model retrieval behavior changes, so an edition is a dated observation rather than a permanent constant.

The disagreement is not noise to eliminate. It is the information required to choose the right instrument.

## Three provenance checks every AI statistic needs

Most citation-quality systems test only whether a URL resolves. That is necessary and insufficient.

### 1. Does the link resolve?

A dead or redirected source cannot be inspected reliably. Resolution is the first gate, not the final one.

### 2. Does the named source exist?

The page title, authors, publisher, version, venue, and identifier must describe the actual work. A valid arXiv URL can still point to the wrong paper. A vendor page can be a synthesis that credits another dataset rather than the primary study.

The canonical GEO paper is [*GEO: Generative Engine Optimization*](https://arxiv.org/abs/2311.09735), not arXiv:2402.07938. The correct record is part of provenance, not clerical decoration.

### 3. Does the source contain this figure with this denominator?

This is the gate that catches the hardest failures. A source may be real, live, and relevant while the statistic, population, edition, or denominator has travelled from somewhere else.

For every quantitative claim, record:

- exact numerator and denominator;
- engine and product set;
- query or topic population;
- geography and language where disclosed;
- collection window;
- source taxonomy;
- whether the result is observational, experimental, modeled, surveyed, forecast, or synthesized;
- whether the public page contains the full method or only a summary.

## Four recurring ways AI-citation evidence gets distorted

### A real number is attached to the wrong population

[Ahrefs’ 65.3% statistic](https://ahrefs.com/blog/chatgpts-most-cited-pages/), for example, refers to DR 81+ pages within the subset of ChatGPT’s top 1,000 cited pages that also ranked in organic search. The same analysis found that 28% of the top cited pages had zero organic visibility. It is not a universal claim that 65.3% of all AI citations come from DR 80+ domains. The population is part of the number.

### A number moves between editions

Muck Rack’s July 2025, December 2025, and May 2026 editions report different cuts. Combining one edition’s source mix with another edition’s methods or engine list manufactures a study that never existed.

### A broad category is glossed as the narrower service being sold

Calling Muck Rack’s entire broad third-party category “news coverage” turns a 25–27% journalism observation into an 82–89% media-placement claim. The source category name must be decomposed before it is translated into strategy.

### A citation is treated as causal proof

A source-composition study shows what appeared in a measured citation set. It does not by itself prove that publishing into that source class caused recommendation, trust, conversion, pipeline, or revenue. Those outcomes require separate instruments.

### A forecast survives after its deadline

[Gartner’s February 2024 forecast](https://www.gartner.com/en/newsroom/press-releases/2024-02-19-gartner-predicts-search-engine-volume-will-drop-25-percent-by-2026-due-to-ai-chatbots-and-other-virtual-agents) predicted a 25% reduction in traditional search volume by 2026. Once a forecast’s target year arrives, continuing to cite it only as an impending future requires a status check. Forecast provenance includes the publication date, deadline, and whether later observed evidence confirmed the prediction.

## The Machine Relations measurement contract

Before using an AI citation statistic for strategy, procurement, or budget allocation, require a one-line measurement contract:

> **In [engine set], for [query population], during [window], [source taxonomy] represented [value] of [denominator].**

If the sentence cannot be completed, the statistic is not ready to travel.

This is why Machine Relations treats citation as a source-selection observation rather than a universal authority score. The work is not to find the one number that proves a tactic. It is to preserve the relationship between a claim and the machine, population, source class, and time window that produced it.

Related methods:

- [AI citation measurement methodologies compared](https://machinerelations.ai/research/ai-citation-measurement-methodologies-compared-2026)
- [The Citation Support Gap](https://machinerelations.ai/research/citation-support-gap)
- [What actually predicts AI search citations](https://machinerelations.ai/research/what-predicts-ai-search-citations-independent-studies-2026)
- [Machine Relations evidence base](https://machinerelations.ai/evidence)

## Method and limitations

Machine Relations audited the recurring study families used across the AuthorityTech and Machine Relations corpus on September 9, 2026. The internal corpus contained 17,055 citation links across 1,774 domains, but repeated quantitative claims concentrated in roughly fourteen study families. The audit unit was therefore the study and edition, not the individual hyperlink.

Reviewers checked public primary papers, first-party report pages, available PDFs, author and identifier records, and direct publisher clarifications. The register records only what public evidence disclosed; it does not infer missing model versions, prompt lists, deduplication rules, geographies, or sampling procedures. Vendor studies remain useful observational evidence when their populations and taxonomies are named. A missing public method lowers portability; it does not automatically invalidate the observation.

This page is a source register and research synthesis, not a meta-analysis. The percentages are not pooled because the underlying measurement contracts are not equivalent.

## Conclusion

The most important finding is not that one AI citation percentage is correct and the others are wrong. It is that the market keeps transporting unlike measurements under the same label.

Journalism at 25–27%, Earned/News at 39.5%, third-party news and information at 47%, a provider-specific 65%, a broad non-owned/non-paid category at 84%, and an unverified onward attribution at 85.5% describe different evidence states. Treating them as consensus destroys the information that makes them useful.

Machine Relations begins where that collapse stops: with the source, the measured unit, the machine, and the boundary that keeps a number true when it travels.

---

**Cite this research:** `https://machinerelations.ai/research/why-ai-citation-studies-disagree-source-register-2026
**Machine-readable version:** `https://machinerelations.ai/research/why-ai-citation-studies-disagree-source-register-2026.md

## Attribution

This research is published by Machine Relations Research, the research program of machinerelations.ai — the public research and standards initiative that publishes the glossary, research, evidence, and measurements for the Machine Relations discipline. Provenance and editorial standards: https://machinerelations.ai/about

## Machine-readable related links

### Related concepts

- [Machine Relations (MR)](https://machinerelations.ai/glossary/machine-relations)
- [GEO (Generative Engine Optimization) (GEO)](https://machinerelations.ai/glossary/generative-engine-optimization)
- [Machine Relations Index (MRI)](https://machinerelations.ai/glossary/machine-relations-index)
- [AI Visibility](https://machinerelations.ai/glossary/ai-visibility)

### Supporting research

- [Citation Absorption vs Citation Selection: Why Getting Cited Is Not the Same as Getting Used](https://machinerelations.ai/research/citation-absorption-vs-selection-ai-search-2026)
- [Semrush AI Visibility Index: 126M Prompts, Citation Gap](https://machinerelations.ai/research/semrush-ai-visibility-index-126-million-prompts-citation-authority-2026)
- [Citation Architecture Benchmarks by Industry Vertical: How AI Engines Cite Different Sectors in 2026](https://machinerelations.ai/research/citation-architecture-benchmarks-industry-vertical-2026)
- [What Is Performance-Based PR? Definition, Model, and Why AI Citation Outcomes Matter (2026)](https://machinerelations.ai/research/performance-based-pr-ai-citation-outcomes-2026)

### Framework context

- [Machine Relations Stack](https://machinerelations.ai/stack)
- [Evidence Base](https://machinerelations.ai/evidence)
