Two AI answer engines support the claims they cite at almost the same rate. Of the citations that can be pinned to a discrete claim and read back, ChatGPT supports 54.8% and Perplexity supports 50.1%. Put those two numbers on a slide and the engines look interchangeable.
They are not. The half that fails, fails in opposite directions.
When a ChatGPT citation does not support its claim, the most common reason is that the page is about something else, or is a real article that supports nothing in the claim at all: 31.3% of its resolvable citations. When a Perplexity citation does not support its claim, the most common reason is that the page is on-topic and partly right and the answer reached past it: 32.7% amplified, 2.85 times ChatGPT's rate, while its wrong-page share is 15.6%, half of ChatGPT's.
That distinction is the whole practical content of a citation. A count of citations cannot see it, and a support rate averages it away.
What was measured, and what the numbers are not #
Every figure below comes from Answer-Source Fidelity, a receipt-bound instrument that takes one AI answer, splits it into atomic claims, ties each citation to the specific claim it anchors, fetches the cited page, and has two graders from different model families label the (claim, page) pair independently, with a third family adjudicating disagreements blind. The live aggregate is published nightly at /measurement/answer-source-fidelity/data.json; this reading is aggregate 90f6711d…, generated 2026-09-19T11:34:53Z, cycle 17.
The label set has five values, and the definitions are the instrument's own:
| Label | The cited page… |
|---|---|
| Supported | contains text supporting the specific claim |
| Amplified | is on-topic and partly supports it, but the claim overstates, generalizes, or adds specifics the page does not contain |
| Contradicted | directly contradicts the claim, including an altered number, entity or date |
| Misattributed | is real but about a different subject; it does not contain this claim |
| Fabricated | is a real, on-topic-ish article that still supports nothing in the claim — the citation is decorative |
A sixth operational label, Unreadable, keeps a page nobody could fetch from being scored as a fabrication.
Three constraints govern how these numbers may be stated, and they are the instrument's, not editorial preference. Rates are reported per engine over the resolvable set only, with resolvable coverage stated beside them. Rates below the evidence floor — fewer than 100 resolvable citations or under 20% coverage — are withheld rather than estimated. And cohorts with different measurement semantics are reported on separate tracks and never pooled, because a blended number is dominated by whichever cohort happened to be most resolvable.
Two engines clear the floor. Gemini (392 resolvable of 4,349, 9.0% coverage) and Claude (88 of 1,361, 6.5%) do not, and carry no rate here. Withholding is not hiding: below the floor the honest output is "not yet."
Finding 1 — Two engines, one rate, opposite failure modes #
Legacy baseline cohort, 10,066 graded events from the frozen corpus, cited pages fetched at grade time.
| ChatGPT | Perplexity | |
|---|---|---|
| Resolvable citations | 1,255 of 1,630 | 1,298 of 2,726 |
| Resolvable coverage | 77.0% | 47.6% |
| Supported | 54.8% (688) | 50.1% (650) |
| Amplified | 11.5% (144) | 32.7% (425) |
| Misattributed | 13.1% (164) | 2.3% (30) |
| Fabricated | 18.3% (229) | 13.3% (173) |
| Contradicted | 2.4% (30) | 1.5% (20) |
Read the failure rows, not the headline row.
Perplexity's characteristic failure is overreach on a defensible source. A third of its resolvable citations land on a page that is genuinely about the subject and genuinely says something adjacent — and the sentence beside it claims more than the page does. Its wrong-page rate is 2.3% misattributed, which is one-fifth of ChatGPT's.
ChatGPT's characteristic failure is the wrong page. Misattributed and Fabricated together take 393 of its 1,255 resolvable citations — 31.3%. Nearly a third of the time, a ChatGPT citation sits next to a claim its page does not support at all, either because the page is about a different subject or because it is an on-topic article that happens to support nothing in the sentence it was attached to.
For a publisher this is not a tie. A Perplexity citation is comparatively strong evidence that your page was read as relevant to the question, and the risk it carries is that your page becomes the evidence for a stronger claim than you made. A ChatGPT citation is much weaker evidence of topical fit: roughly one in three is decorative or plain wrong-subject, and no amount of it changes what the page says.
Finding 2 — The shape holds on an independent cohort with a different clock #
The obvious objection is timing. The legacy baseline fetches cited pages at grade time, days after the answer was produced, so it measures whether the page supports the claim now. Pages change. Maybe ChatGPT's wrong-page share is really link rot.
It is not. A second cohort runs nightly on a different fetch basis: the cited page is snapshotted minutes after the answer is produced and grading consumes only that snapshot. It has run 17 cycles since 2026-09-02, and it holds ChatGPT only.
| Label | Legacy baseline (grade-time fetch) | Collection-time series (snapshot at citation) | Delta |
|---|---|---|---|
| Supported | 54.8% | 52.0% | −2.8 |
| Fabricated | 18.3% | 20.1% | +1.8 |
| Misattributed | 13.1% | 13.8% | +0.7 |
| Amplified | 11.5% | 12.0% | +0.5 |
| Contradicted | 2.4% | 2.1% | −0.3 |
| Resolvable / total | 1,255 / 1,630 | 4,013 / 5,459 | — |
| Resolvable coverage | 77.0% | 73.5% | — |
Two cohorts, two fetch clocks, two denominators, 5,268 resolvable citations between them — and every label lands within three points. These are separate tracks and are not pooled into one number. Their agreement is the point: ChatGPT's failure profile is a property of which page it picks, not of when the page is read.
Finding 3 — The question shape decides which way the citation fails #
The failure mode is not fixed per engine either. Within the collection-time series, splitting by the shape of the buyer's question moves it by more than 20 points.
Every row below clears the evidence floor (443 or more resolvable citations, 71.8% coverage or better).
| Question shape | Resolvable | Coverage | Supported | Fabricated | Misattributed | Amplified |
|---|---|---|---|---|---|---|
| "is X worth it" | 462 | 76.9% | 64.3% | 16.7% | 3.9% | 11.9% |
| "X vs Y" | 443 | 74.0% | 62.5% | 17.8% | 5.4% | 12.9% |
| "how do I choose" | 571 | 73.7% | 49.9% | 33.1% | 6.1% | 9.1% |
| "I have this problem" | 570 | 73.6% | 48.9% | 33.2% | 5.8% | 11.1% |
| "top list" | 935 | 71.8% | 48.4% | 12.5% | 25.5% | 12.3% |
| "best X" | 1,032 | 73.3% | 47.9% | 15.2% | 19.9% | 13.5% |
Three regimes, and each one is a different job for the engine.
Decision questions cite best. "Is X worth it" and "X vs Y" support at 62–64%, with misattribution under 6%. These questions have a discrete answerable claim, and there is usually a page that makes it.
Guidance questions produce decorative citations. Ask "how do I choose" or describe a problem, and a third of the resolvable citations are Fabricated — a real, on-topic article attached to a sentence it does not support. The engine is writing the guidance itself and stapling a plausible source to it. Misattribution collapses to about 6% because the page it grabs is on-topic; it just is not evidence.
List questions produce the wrong page. "Best X" and "top list" carry the highest misattribution of any shape, 19.9% and 25.5%, more than four times the decision shapes. Listicles cite listicles, and the one that gets cited is frequently a list about a neighbouring subject.
This is the fidelity counterpart to a pattern already measured on the source side: the question shape decides which source wins. It also decides whether the citation means anything.
The leg a publisher actually controls #
One failure class is not the engine's judgment at all. In the collection-time series, 1,063 of 5,459 cited pages — 19.5% — could not be read back at all, after a fetch ladder that tries a direct request, then a residential proxy, then a rendering service, and only terminalizes an event once repeated attempts across nights hit a retry cap.
That splits two ways: 675 where every rung of the ladder failed, and 388 where something was fetched and both graders agreed it was not an article — a bot wall, a consent shell, a login or paywall gate, an "enable JavaScript" stub, or pure navigation boilerplate.
A page an independent client cannot read is a page whose claim cannot be verified by anyone downstream of the answer — a grader, a second engine, or a reader who clicks. It is also the one number on this page that a publisher can move this week, without persuading an engine of anything.
Bounds, and the one that cuts against the headline #
The external validation is a binary one. The instrument was checked against AttrScore, a public human-annotated attribution benchmark (paper; ACL version; code), pinned at dataset revision 467dcdd2, on a balanced three-way sample: 55 of 60 correct, 91.7% agreement, Wilson lower bound 81.9%. That certificate covers the supported / not-supported call. It does not separately validate each of the five sub-labels, and the benchmark's own per-class figures are explicitly diagnostic, never a publication gate.
The per-class pattern matters here, and it works against this page's own framing. On AttrScore's extrapolatory class — where the claim reaches beyond its source — the committee agreed on only 6 of 18, pushing 11 to not-attributable and letting 1 through as attributable. In plain terms: where a human annotator says "overreach," this instrument usually says "unsupported." So Amplified is, if anything, under-counted relative to a human, and the harsher wrong-page bucket is over-counted.
That bias is why the contrast is the finding and the absolute splits are directional. Both engines in Finding 1 sit in one cohort, graded by one committee under one policy, so a shared strictness cannot manufacture an opposite ordering between them. It could move ChatGPT's 31.3% wrong-page share by some points. It cannot turn Perplexity's 32.7% amplification into ChatGPT's 11.5%.
The rest of the bounds, stated rather than buried: inter-model raw agreement is 72.2% and Cohen's κ is 0.647; control fixtures run 30 of 30; the sealed calibration certificate scores 83 of 91 (91.2%, Wilson lower bound 83.6%); the primary-excluded holdout is not available (n = 0) and is a diagnostic, not a gate. Coverage differs sharply between the two engines in Finding 1 — 77.0% against 47.6% — and coverage is itself a property of answer style, so the two rates describe differently-sized slices of each engine's output. Gemini and Claude are unmeasured at claim level, not clean. The corpus is a fixed basket of English-language commercial questions across twelve verticals; it cannot see a market it was not pointed at.
What would change this #
A falsifiable claim needs a falsifier. This one has three. If ChatGPT's misattribution and Perplexity's amplification converge over the coming cycles, the split is a transient of two product generations rather than two retrieval architectures. If the collection-time series diverges from the legacy baseline by more than a few points on any label as it accumulates, the grade-time bound matters more than this page allows. And if Gemini or Claude clear the coverage floor and land in a third pattern, the two-regime reading here is incomplete.
The series regenerates nightly and is published in machine-readable form. Every number above is checkable against it.
Methodology and sources #
The methodology is the published Answer-Source Fidelity contract: atomic (claim, source) pairs; always-dual grading by two model families with blind third-family adjudication on disagreement; an adversarial refuter attempting to overturn every accepted label; abstention rather than a forced label; deterministic control fixtures seeded through each run so that a mid-run control failure halts it; and a content-hashed receipt chain from corpus to published number. Cohorts partition mechanically by measurement semantics — fetch basis, prompt hash, model pins, claim-map version, label set, adjudication policy, denominator and evidence policy — and identical semantics pool automatically while any semantic change opens a new track.
Source of every figure: the nightly aggregate at /measurement/answer-source-fidelity/data.json, aggregate_sha256 beginning 90f6711d, generated 2026-09-19T11:34:53Z, 17 cycles and 15,525 denominator events since 2026-09-02. Cohort totals, coverage and the five-label distribution over the resolvable denominator are read directly from that artifact; the question-shape split is computed over the same cohort's event-level grade index, 5,459 events, and reconciles exactly to the published cohort totals (4,013 graded, 1,063 unreadable, 383 uncertain). The binary support rates and coverage restated in Finding 1 are the same figures carried on the canonical report, The Citation Support Gap, and on the per-engine support rate brief.
External benchmark and prior work: AttrScore / AttrEval-GenSearch at pinned dataset revision 467dcdd2 (paper, ACL Anthology, code). Two prior public measurements are adjacent but measure different objects, and neither is restated as this instrument's result. Liu, Zhang and Liang's human audit of four generative search engines found "a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence" — sentence-level support across a 2023 engine generation, not claim-level support over a resolvable denominator. The Tow Center's 2025 audit of eight AI search tools found they "provided incorrect answers to more than 60 percent of queries", measuring whether a chatbot could identify the source article behind an excerpt — a retrieval-identification task, not whether the citations it emitted support the claims beside them.
Engine behaviour referenced: OpenAI's web search tool documentation, Perplexity's search API guides, Google's AI Mode announcement and AI features documentation, Google Search help on AI experiences, and Anthropic's Citations API and its developer documentation. Definitional background on a citation that points at nothing: hallucinated citation.
Reproducibility: pinned model identifiers, prompt hashes, code hashes, and a content-hashed receipt chain; re-running aggregation over identical frozen matter is idempotent and reproduces aggregate_sha256.