Research

When an AI Citation Index Is Wrong: One Misclassified Field, 57 Corrected Pages

A source-classification defect in one Machine Relations Index release put a wire distributor inside the editorial class. Correcting it took 57 live pages across six domains in a single day. This is the correction record, the arithmetic, and the standard we are holding ourselves to.

Published Machine Relations Research
Reference
TopicsMeasurementCitationsAI SearchMethodologyCorrections

When a published measurement turns out to be wrong, the number does not live on one page. It lives on every page that quoted it. On September 20, 2026 a single misclassified field in the Machine Relations Index reached 57 live pages across six domains, and correcting it took five hours. This is the record of what was wrong, what it cost, and what still stands.

What was wrong #

The Index assigns every cited domain a source_role — editorial publication, wire and press-release distribution, community platform, vendor-owned, and five others. In release mri_score_v2.0+2026-09-20+47973f373a20, generated September 20, 2026 over the window May 10 to September 20, 2026, that field files six domains as editorial publications that are not edited newsrooms: medium.com, amazon.com, prnewswire.com, apple.com, wordpress.com and blogspot.com.

The clearest case is the wire one. PR Newswire's public profile in that release carries the source role Editorial publication and a standing of #8 of 1,230 classified editorial publications, with a citation rate of 1.13% — cited in 182 of 16,039 monitored answer runs, on 75 of 127 days, across all six observed engines, at Confidence B. Its two direct competitors are filed correctly: businesswire.com carries the role Wire and press-release distribution at a 0.61% citation rate, cited in 98 of 16,039 runs, and globenewswire.com carries the same role.

So the largest wire distributor in the index sat inside the editorial set, ranked eighth of it, while the two companies it competes with sat outside. At the top of the same class, medium.com — an open publishing platform where anyone can post — holds #1 of 1,230 at a 5.04% citation rate, cited in 808 of 16,039 runs at Confidence A, while substack.com, beehiiv.com and dev.to, the same hosted open-publishing product, are filed as community platforms.

Together those six domains hold 1,358 of the editorial class's 13,525 cited runs, which is 10.04% of the class, and they include its rank-1, rank-7 and rank-8 positions. The full arithmetic is published separately at jaxonparrott.com, and the Amazon case is worked through at paralabs.ai.

What the defect does not touch #

This is the part a correction usually leaves out, and it is the part a reader most needs.

Citation rate, cited runs, days cited, the set of engines a domain was observed on and confidence are per-domain observations over a fixed set of monitored answer runs. A domain either appeared as a source in a run or it did not. No classification decision touches that, so every one of those figures stands exactly as published, in this release and in the ones before it.

What the defect does touch is anything computed across a source-role class: class share, class rank, and every editorial-versus-wire or editorial-versus-community comparison. Those were wrong wherever we published them, and they are what we corrected.

What it cost to correct #

The load-bearing casualty was our own wire-versus-editorial number. We had published the wire and press-release class as holding 0.21% of classified citation, and that figure was computed over a nine-domain class that excluded the largest wire domain in the index. Counted where it belongs, the class is ten domains holding 414 of 112,518 classified cited-domain observations — 0.37%, nearly double what we published.

The finding survived the correction. At 414 observations the wire and press-release class is still the smallest classified citation pool in the release by a wide margin. But the number we had been repeating was understated, and we had repeated it a great deal.

On September 20, 2026 the correction ran across every property we publish:

Property Pages carrying a correction
authoritytech.io 24
jaxonparrott.com 21
machinerelations.ai 8
christianlehman.com 2
paralabs.ai 1
paralax.ai 1
Total 57

Those 57 pages are every live page that carries either the restated 0.37% figure or the canonical classification-limit note placed beside an affected class figure, read from the published main branch of all six site repositories. The wider day's activity was 124 publisher commits producing 62 distinct updated pages and 9 new ones; the classification note itself was introduced between 15:58 and 21:00 UTC.

One field. Five hours. Fifty-seven pages.

What measurement institutions do about this #

Every serious measurement institution has already answered the question of what you owe a published number after you publish it. The answers are old and they are consistent.

Media measurement. The Media Rating Council's Minimum Standards for Media Rating Research, the assessment basis for accrediting audience measurement services, states it as Disclosure Standard B.2 in the December 2011 revision: "Each rating report shall point out changes in, or deviations from, the standard operating procedures of the rating service which may exert a significant effect on the reported results. This notification shall indicate the estimated magnitude of the effect. The notice shall go to subscribers in advance as well as being prominently displayed in the report itself." Not a correction when asked. A disclosure with a magnitude attached, in the report.

Search measurement. Google maintains a public data anomalies log for Search Console, recording known issues that affect reported data. Its entries include a logging error that prevented Search Console from accurately reporting impressions from May 13, 2025 until April 27, 2026 — nearly a year of a metric the entire SEO industry reports to boards, disclosed in a dated public list.

Scientific publishing. The ICMJE recommendations carry a section on corrections and version control. Nature Portfolio's correction and retraction policy defines distinct amendment types and links every one bidirectionally to the original article. Crossmark exists so a reader holding a document can check whether it is still the current version, and Retraction Watch maintains a public database of what was withdrawn. arXiv keeps every prior version addressable rather than overwriting it.

Official statistics. The Bureau of Economic Analysis publishes each quarter's GDP in successive vintages — an advance estimate, then a second estimate, then later revisions — and maintains revision information and a vintage history alongside the current figure, stating plainly that previously published estimates have since been revised. The discipline of metrology writes the same idea into its foundations: the JCGM guides published through the BIPM treat a measurement result as incomplete without a stated uncertainty.

Digital advertising. IAB Tech Lab maintains versioned measurement specifications precisely so that a change in what is counted is visible as a change, rather than appearing as a change in the thing being counted.

The pattern across all of them is the same, and it is not "be accurate." It is: publish what changed, say how big the change was, and leave the old number addressable. AI citation measurement is roughly two years old and inherits none of this by default. It is the newest measurement discipline in marketing and the only one currently making load-bearing claims with no shared convention for what happens when those claims move.

The correction standard we hold #

Machine Relations is the discipline of earning AI-engine citations, and an index that cannot correct backwards is not an instrument. Five rules, which this page is the first application of.

  1. Every published figure carries its release identifier and observation window. A number without a traceable artifact behind it cannot be corrected later, because nobody can tell which run produced it.
  2. A correction runs to every page carrying the affected figure, not only to the page where the error was found. The unit of correction is the claim, not the document.
  3. The old number stays visible next to the new one. A claim that is quietly deleted teaches a reader nothing and hides the size of the error.
  4. What the defect does not affect is stated as explicitly as what it does. Over-withdrawing is its own failure: it destroys sound measurement to look contrite.
  5. The correction record is public and dated. This page is it, and it will be appended to rather than replaced.

Rule 3 has a consequence worth naming, because we learned it the hard way inside this very sweep. A remedied page still contains the wrong figure — that is the point of leaving it visible — so a text search for the bad number finds corrected pages and uncorrected ones alike. A correction sweep that greps for the claim and stops there will report failures that are fixes. Read the survivors live, next to their remedy, or the audit is measuring itself.

Correction record #

Date Release What was wrong Restated to Pages
2026-09-20 mri_score_v2.0+2026-09-20+47973f373a20 source_role files prnewswire.com, medium.com, amazon.com, apple.com, wordpress.com and blogspot.com as editorial publications; class share, class rank and editorial-versus-wire comparisons computed over a contaminated set Wire and press-release class share restated from 0.21% to 0.37%, 414 of 112,518 classified cited-domain observations; a classification-limit note placed beside every affected class figure; per-domain rates, cited runs, days cited, observed engines and confidence unaffected 57

The classifier defect itself is not yet fixed in the release. Until it is, read the rate, not the class rank.

Sources and method #

Index figures were read live on September 20, 2026 from the public per-domain profiles of release mri_score_v2.0+2026-09-20+47973f373a20, methodology mri_score_v2.0, window May 10 to September 20, 2026, 16,039 monitored answer runs across ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews and Perplexity: prnewswire.com, medium.com and businesswire.com. Every domain in the index has a public profile under machinerelations.ai/index/domains.

Page counts were computed from the published main branch of the six site repositories that host every property named here, by listing the files containing either the restated 0.37% figure or the canonical classification-limit note, deduplicating, and grouping by repository: 24, 21, 8, 2, 1 and 1, totalling 57 files, all of them published content. Commit and update counts for September 20, 2026 UTC were taken from the same branches.

External standards and policies were read at each publisher's own page on September 20, 2026, and each is linked above at the URL read: the Media Rating Council's Minimum Standards for Media Rating Research (December 2011 revision, quoted verbatim from the PDF published at mediaratingcouncil.org), Google's Search Console data anomalies log, the ICMJE recommendations, Nature Portfolio's correction and retraction policy, Crossref's Crossmark, the Retraction Watch database guide, arXiv's submission and versioning help, the Bureau of Economic Analysis GDP page carrying revision information and vintage history, the JCGM publications hosted by the BIPM, and IAB Tech Lab's standards index. No third-party roundup was used as evidence for what any of these organisations publishes.

The corrections described here were published through our own guarded publisher and verified at the live URL. The corrected origin page for the concentration finding is how concentrated are AI citations, and the corrected wire-versus-editorial argument is at authoritytech.io.