# Same Support Rate, Opposite Failure Modes: What ChatGPT and Perplexity Do When a Citation Doesn't Hold

ChatGPT and Perplexity support their claims at almost the same rate — 54.8% and 50.1% of resolvable citations. The rate hides the difference that matters. When a ChatGPT citation fails, it usually points at the wrong page (31.3% misattributed or decorative). When a Perplexity citation fails, it usually points at a reasonable page and overstates it (32.7% amplified, 2.85x ChatGPT). A five-label failure-mode split from the Answer-Source Fidelity instrument, with resolvable coverage stated beside every rate.

Canonical URL: https://machinerelations.ai/research/ai-citation-failure-modes-chatgpt-perplexity-2026
Published: 2026-09-19
Research type: Study

## Source Body

Two AI answer engines support the claims they cite at almost the same rate. Of the citations that can be pinned to a discrete claim and read back, ChatGPT supports **54.8%** and Perplexity supports **50.1%**. Put those two numbers on a slide and the engines look interchangeable.

They are not. The half that fails, fails in opposite directions.

When a ChatGPT citation does not support its claim, the most common reason is that the page is about something else, or is a real article that supports nothing in the claim at all: **31.3% of its resolvable citations**. When a Perplexity citation does not support its claim, the most common reason is that the page is on-topic and partly right and the answer reached past it: **32.7% amplified**, 2.85 times ChatGPT's rate, while its wrong-page share is **15.6%**, half of ChatGPT's.

That distinction is the whole practical content of a citation. A count of citations cannot see it, and a support rate averages it away.

## What was measured, and what the numbers are not

Every figure below comes from [Answer-Source Fidelity](https://machinerelations.ai/research/citation-support-gap), a receipt-bound instrument that takes one AI answer, splits it into atomic claims, ties each citation to the specific claim it anchors, fetches the cited page, and has two graders from different model families label the (claim, page) pair independently, with a third family adjudicating disagreements blind. The live aggregate is published nightly at [`/measurement/answer-source-fidelity/data.json`](https://machinerelations.ai/measurement/answer-source-fidelity/data.json); this reading is aggregate `90f6711d…`, generated 2026-09-19T11:34:53Z, cycle 17.

The label set has five values, and the definitions are the instrument's own:

| Label | The cited page… |
| --- | --- |
| Supported | contains text supporting the specific claim |
| Amplified | is on-topic and partly supports it, but the claim overstates, generalizes, or adds specifics the page does not contain |
| Contradicted | directly contradicts the claim, including an altered number, entity or date |
| Misattributed | is real but about a different subject; it does not contain this claim |
| Fabricated | is a real, on-topic-ish article that still supports nothing in the claim — the citation is decorative |

A sixth operational label, Unreadable, keeps a page nobody could fetch from being scored as a fabrication.

Three constraints govern how these numbers may be stated, and they are the instrument's, not editorial preference. Rates are reported per engine over the **resolvable** set only, with resolvable coverage stated beside them. Rates below the evidence floor — fewer than 100 resolvable citations or under 20% coverage — are withheld rather than estimated. And cohorts with different measurement semantics are reported on separate tracks and never pooled, because a blended number is dominated by whichever cohort happened to be most resolvable.

Two engines clear the floor. Gemini (392 resolvable of 4,349, 9.0% coverage) and Claude (88 of 1,361, 6.5%) do not, and carry no rate here. Withholding is not hiding: below the floor the honest output is "not yet."

## Finding 1 — Two engines, one rate, opposite failure modes

Legacy baseline cohort, 10,066 graded events from the frozen corpus, cited pages fetched at grade time.

| | ChatGPT | Perplexity |
| --- | ---: | ---: |
| Resolvable citations | 1,255 of 1,630 | 1,298 of 2,726 |
| Resolvable coverage | 77.0% | 47.6% |
| **Supported** | **54.8%** (688) | **50.1%** (650) |
| Amplified | 11.5% (144) | **32.7%** (425) |
| Misattributed | 13.1% (164) | 2.3% (30) |
| Fabricated | 18.3% (229) | 13.3% (173) |
| Contradicted | 2.4% (30) | 1.5% (20) |

Read the failure rows, not the headline row.

**Perplexity's characteristic failure is overreach on a defensible source.** A third of its resolvable citations land on a page that is genuinely about the subject and genuinely says something adjacent — and the sentence beside it claims more than the page does. Its wrong-page rate is 2.3% misattributed, which is one-fifth of ChatGPT's.

**ChatGPT's characteristic failure is the wrong page.** Misattributed and Fabricated together take 393 of its 1,255 resolvable citations — 31.3%. Nearly a third of the time, a ChatGPT citation sits next to a claim its page does not support at all, either because the page is about a different subject or because it is an on-topic article that happens to support nothing in the sentence it was attached to.

For a publisher this is not a tie. A Perplexity citation is comparatively strong evidence that your page was read as relevant to the question, and the risk it carries is that your page becomes the evidence for a stronger claim than you made. A ChatGPT citation is much weaker evidence of topical fit: roughly one in three is decorative or plain wrong-subject, and no amount of it changes what the page says.

## Finding 2 — The shape holds on an independent cohort with a different clock

The obvious objection is timing. The legacy baseline fetches cited pages at grade time, days after the answer was produced, so it measures whether the page supports the claim *now*. Pages change. Maybe ChatGPT's wrong-page share is really link rot.

It is not. A second cohort runs nightly on a different fetch basis: the cited page is snapshotted minutes after the answer is produced and grading consumes only that snapshot. It has run 17 cycles since 2026-09-02, and it holds ChatGPT only.

| Label | Legacy baseline (grade-time fetch) | Collection-time series (snapshot at citation) | Delta |
| --- | ---: | ---: | ---: |
| Supported | 54.8% | 52.0% | −2.8 |
| Fabricated | 18.3% | 20.1% | +1.8 |
| Misattributed | 13.1% | 13.8% | +0.7 |
| Amplified | 11.5% | 12.0% | +0.5 |
| Contradicted | 2.4% | 2.1% | −0.3 |
| Resolvable / total | 1,255 / 1,630 | 4,013 / 5,459 | — |
| Resolvable coverage | 77.0% | 73.5% | — |

Two cohorts, two fetch clocks, two denominators, 5,268 resolvable citations between them — and every label lands within three points. These are separate tracks and are not pooled into one number. Their agreement is the point: ChatGPT's failure profile is a property of which page it picks, not of when the page is read.

## Finding 3 — The question shape decides which way the citation fails

The failure mode is not fixed per engine either. Within the collection-time series, splitting by the shape of the buyer's question moves it by more than 20 points.

Every row below clears the evidence floor (443 or more resolvable citations, 71.8% coverage or better).

| Question shape | Resolvable | Coverage | Supported | Fabricated | Misattributed | Amplified |
| --- | ---: | ---: | ---: | ---: | ---: | ---: |
| "is X worth it" | 462 | 76.9% | **64.3%** | 16.7% | 3.9% | 11.9% |
| "X vs Y" | 443 | 74.0% | **62.5%** | 17.8% | 5.4% | 12.9% |
| "how do I choose" | 571 | 73.7% | 49.9% | **33.1%** | 6.1% | 9.1% |
| "I have this problem" | 570 | 73.6% | 48.9% | **33.2%** | 5.8% | 11.1% |
| "top list" | 935 | 71.8% | 48.4% | 12.5% | **25.5%** | 12.3% |
| "best X" | 1,032 | 73.3% | 47.9% | 15.2% | **19.9%** | 13.5% |

Three regimes, and each one is a different job for the engine.

**Decision questions cite best.** "Is X worth it" and "X vs Y" support at 62–64%, with misattribution under 6%. These questions have a discrete answerable claim, and there is usually a page that makes it.

**Guidance questions produce decorative citations.** Ask "how do I choose" or describe a problem, and a third of the resolvable citations are Fabricated — a real, on-topic article attached to a sentence it does not support. The engine is writing the guidance itself and stapling a plausible source to it. Misattribution collapses to about 6% because the page it grabs *is* on-topic; it just is not evidence.

**List questions produce the wrong page.** "Best X" and "top list" carry the highest misattribution of any shape, 19.9% and 25.5%, more than four times the decision shapes. Listicles cite listicles, and the one that gets cited is frequently a list about a neighbouring subject.

This is the fidelity counterpart to a pattern already measured on the source side: [the question shape decides which source wins](https://authoritytech.io/curated/question-shape-decides-ai-citation-winner-2026). It also decides whether the citation means anything.

## The leg a publisher actually controls

One failure class is not the engine's judgment at all. In the collection-time series, 1,063 of 5,459 cited pages — 19.5% — could not be read back at all, after a fetch ladder that tries a direct request, then a residential proxy, then a rendering service, and only terminalizes an event once repeated attempts across nights hit a retry cap.

That splits two ways: 675 where every rung of the ladder failed, and 388 where something was fetched and both graders agreed it was not an article — a bot wall, a consent shell, a login or paywall gate, an "enable JavaScript" stub, or pure navigation boilerplate.

A page an independent client cannot read is a page whose claim cannot be verified by anyone downstream of the answer — a grader, a second engine, or a reader who clicks. It is also the one number on this page that a publisher can move this week, without persuading an engine of anything.

## Bounds, and the one that cuts against the headline

The external validation is a **binary** one. The instrument was checked against [AttrScore](https://huggingface.co/datasets/osunlp/AttrScore), a public human-annotated attribution benchmark ([paper](https://arxiv.org/abs/2305.06311); [ACL version](https://aclanthology.org/2023.findings-emnlp.307/); [code](https://github.com/OSU-NLP-Group/AttrScore)), pinned at dataset revision `467dcdd2`, on a balanced three-way sample: 55 of 60 correct, 91.7% agreement, Wilson lower bound 81.9%. That certificate covers the supported / not-supported call. It does not separately validate each of the five sub-labels, and the benchmark's own per-class figures are explicitly diagnostic, never a publication gate.

The per-class pattern matters here, and it works against this page's own framing. On AttrScore's *extrapolatory* class — where the claim reaches beyond its source — the committee agreed on only 6 of 18, pushing 11 to not-attributable and letting 1 through as attributable. In plain terms: where a human annotator says "overreach," this instrument usually says "unsupported." So Amplified is, if anything, under-counted relative to a human, and the harsher wrong-page bucket is over-counted.

That bias is why the *contrast* is the finding and the absolute splits are directional. Both engines in Finding 1 sit in one cohort, graded by one committee under one policy, so a shared strictness cannot manufacture an opposite ordering between them. It could move ChatGPT's 31.3% wrong-page share by some points. It cannot turn Perplexity's 32.7% amplification into ChatGPT's 11.5%.

The rest of the bounds, stated rather than buried: inter-model raw agreement is 72.2% and Cohen's κ is 0.647; control fixtures run 30 of 30; the sealed calibration certificate scores 83 of 91 (91.2%, Wilson lower bound 83.6%); the primary-excluded holdout is not available (n = 0) and is a diagnostic, not a gate. Coverage differs sharply between the two engines in Finding 1 — 77.0% against 47.6% — and coverage is itself a property of answer style, so the two rates describe differently-sized slices of each engine's output. Gemini and Claude are unmeasured at claim level, not clean. The corpus is a fixed basket of English-language commercial questions across twelve verticals; it cannot see a market it was not pointed at.

## What would change this

A falsifiable claim needs a falsifier. This one has three. If ChatGPT's misattribution and Perplexity's amplification converge over the coming cycles, the split is a transient of two product generations rather than two retrieval architectures. If the collection-time series diverges from the legacy baseline by more than a few points on any label as it accumulates, the grade-time bound matters more than this page allows. And if Gemini or Claude clear the coverage floor and land in a third pattern, the two-regime reading here is incomplete.

The series regenerates nightly and is published in machine-readable form. Every number above is checkable against it.

## Methodology and sources

The methodology is the published Answer-Source Fidelity contract: atomic (claim, source) pairs; always-dual grading by two model families with blind third-family adjudication on disagreement; an adversarial refuter attempting to overturn every accepted label; abstention rather than a forced label; deterministic control fixtures seeded through each run so that a mid-run control failure halts it; and a content-hashed receipt chain from corpus to published number. Cohorts partition mechanically by measurement semantics — fetch basis, prompt hash, model pins, claim-map version, label set, adjudication policy, denominator and evidence policy — and identical semantics pool automatically while any semantic change opens a new track.

Source of every figure: the nightly aggregate at [`/measurement/answer-source-fidelity/data.json`](https://machinerelations.ai/measurement/answer-source-fidelity/data.json), `aggregate_sha256` beginning `90f6711d`, generated 2026-09-19T11:34:53Z, 17 cycles and 15,525 denominator events since 2026-09-02. Cohort totals, coverage and the five-label distribution over the resolvable denominator are read directly from that artifact; the question-shape split is computed over the same cohort's event-level grade index, 5,459 events, and reconciles exactly to the published cohort totals (4,013 graded, 1,063 unreadable, 383 uncertain). The binary support rates and coverage restated in Finding 1 are the same figures carried on the canonical report, [The Citation Support Gap](https://machinerelations.ai/research/citation-support-gap), and on [the per-engine support rate brief](https://paralax.ai/blog/ai-citation-support-rate-per-engine-2026).

External benchmark and prior work: [AttrScore / AttrEval-GenSearch](https://huggingface.co/datasets/osunlp/AttrScore) at pinned dataset revision `467dcdd2` ([paper](https://arxiv.org/abs/2305.06311), [ACL Anthology](https://aclanthology.org/2023.findings-emnlp.307/), [code](https://github.com/OSU-NLP-Group/AttrScore)). Two prior public measurements are adjacent but measure different objects, and neither is restated as this instrument's result. Liu, Zhang and Liang's human audit of four generative search engines found ["a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence"](https://arxiv.org/abs/2304.09848) — sentence-level support across a 2023 engine generation, not claim-level support over a resolvable denominator. The Tow Center's 2025 audit of eight AI search tools found they ["provided incorrect answers to more than 60 percent of queries"](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php), measuring whether a chatbot could identify the source article behind an excerpt — a retrieval-identification task, not whether the citations it emitted support the claims beside them.

Engine behaviour referenced: [OpenAI's web search tool documentation](https://platform.openai.com/docs/guides/tools-web-search), [Perplexity's search API guides](https://docs.perplexity.ai/guides/search-domain-filters), [Google's AI Mode announcement](https://blog.google/products/search/ai-mode-search/) and [AI features documentation](https://developers.google.com/search/docs/appearance/ai-features), [Google Search help on AI experiences](https://support.google.com/websearch/answer/13572151), and Anthropic's [Citations API](https://www.anthropic.com/news/introducing-citations-api) and its [developer documentation](https://docs.anthropic.com/en/docs/build-with-claude/citations). Definitional background on a citation that points at nothing: [hallucinated citation](https://authoritytech.io/glossary/hallucinated-citation).

Reproducibility: pinned model identifiers, prompt hashes, code hashes, and a content-hashed receipt chain; re-running aggregation over identical frozen matter is idempotent and reproduces `aggregate_sha256`.

## Attribution

This research is published by Machine Relations Research, the research program of machinerelations.ai — the public research and standards initiative that publishes the glossary, research, evidence, and measurements for the Machine Relations discipline. Provenance and editorial standards: https://machinerelations.ai/about

## Machine-readable related links

### Related concepts

- [Machine Relations Index (MRI)](https://machinerelations.ai/glossary/machine-relations-index)
- [AI Search Engine](https://machinerelations.ai/glossary/ai-search-engine)
- [MRI Score](https://machinerelations.ai/glossary/mri-score)
- [RAG Citation (RAG)](https://machinerelations.ai/glossary/rag-citation)

### Supporting research

- [The Citation Support Gap: Can AI Citations Be Pinned to the Claims They Anchor?](https://machinerelations.ai/research/citation-support-gap)
- [Brand24 Alternatives in 2026: The AI Citation Gap Every Social Listening Tool Shares](https://machinerelations.ai/research/brand24-alternatives-ai-citation-gap-2026)
- [Citation Absorption vs Citation Selection: Why Getting Cited Is Not the Same as Getting Used](https://machinerelations.ai/research/citation-absorption-vs-selection-ai-search-2026)
- [Cision Alternatives for AI-Era Brand Visibility (2026): What Traditional PR Monitoring Misses](https://machinerelations.ai/research/cision-alternatives-ai-era-2026)

### Framework context

- [Machine Relations Index](https://machinerelations.ai/index)
- [Machine Relations Stack](https://machinerelations.ai/stack)
- [Evidence Base](https://machinerelations.ai/evidence)
