AI citation-rate uncertainty cannot be read from citation-event volume alone. Machine Relations Index rates are denominated by observed answer runs inside a category-question stratum; when those runs are concentrated on seven run dates, the uncertainty question becomes whether runs on the same date behave like independent replications or like clustered observations. (MRI public index)
Why citation events are not independent AI citation-rate replications #
A Machine Relations Index citation event is evidence that a source appeared, but it is not automatically an independent replication of a citation-rate estimate. The September 15, 2026 MRI release reports 122,144 citation events, 15,468 answer runs, 898 monitored prompts, six answer engines, and 122 observed days across the full window from May 10 through September 15. Those are release-scale context fields, not the denominator for every stratum. (MRI release manifest)
Inside the narrower AI Infrastructure strata, the released denominators are much smaller. Reddit appears in 20 of 77 best_x runs, a 25.97% citation rate, across seven run dates. Reddit appears in 17 of 106 how_choose runs, a 16.04% citation rate, also across seven run dates. The public data therefore supports a descriptive comparison between two published strata, but it does not support treating each row as if it had 122 independent daily replications. (MRI public index)
This is different from the September 14 denominator note. That earlier article showed why an overall leaderboard cannot answer a category-specific buying question without changing the eligible answer-run denominator. This note starts after the denominator is selected and asks a second question: how much independent evidence stands behind the rate when the denominator is clustered by date? (AI citation leaderboard category denominator)
Run-date clustering changes the effective AI citation-rate sample size #
Clustered observations can have less effective information than the same number of independent observations. Cochrane's cluster-randomized-trial guidance defines an effective sample size adjustment using a design effect: divide the original sample size by a factor that increases with average cluster size and intracluster correlation. The context is clinical trials, not AI citation measurement, but the statistical caution transfers: correlated observations inside a cluster should not be read as fully independent units. (Cochrane Handbook, Chapter 23)
The same issue appears in longitudinal methods. Liang and Zeger's generalized-estimating-equations paper was built for repeated observations where outcomes within the same subject or time structure can be correlated; its practical lesson for MRI reporting is that variance assumptions belong next to repeated-measurement data, not after the headline rate. (Liang and Zeger, 1986)
For AI citation rates, the likely clustering unit is not a patient, classroom, or household. It is the run date, sometimes crossed with engine, prompt basket, geography, retrieval vendor, or model version. A seven-date stratum can contain 77 or 106 observed runs, but if same-date runs share engine conditions, news context, retrieval freshness, or provider behavior, the independent information can be closer to a smaller effective sample.
A hypothetical clustered-by-day sensitivity example for MRI reporting #
A sensitivity table should show how the same observed citation rate changes under different dependence assumptions without calling those rows live confidence intervals. The table below is deliberately illustrative. It uses the September 15 Reddit AI Infrastructure best_x descriptive rate, 20/77 = 25.97%, then applies the standard design-effect shape 1 + (m - 1)ρ, where m is the average runs per observed date and ρ is a hypothetical same-date correlation. (Cochrane Handbook, Chapter 23)
| Illustrative same-date correlation | Design effect with 77 runs / 7 dates | Effective run count | Approximate standard-error scale | What this row means |
|---|---|---|---|---|
| 0.00 | 1.00 | 77.0 | 4.997 percentage points | Same-date clustering contributes no extra dependence in this toy calculation. |
| 0.05 | 1.50 | 51.3 | 6.120 percentage points | Mild same-date dependence lowers the effective denominator. |
| 0.10 | 2.00 | 38.5 | 7.067 percentage points | Moderate same-date dependence roughly halves the effective run count. |
| 0.20 | 3.00 | 25.7 | 8.655 percentage points | Strong same-date dependence makes the seven-date structure visible in the uncertainty scale. |
The approximate standard-error scale above is computed as sqrt(p(1-p)/n_eff) with p = 20/77. It is not a published MRI confidence interval, because the public MRI release does not expose the run-level clustering structure, engine-by-date cells, or missing-cell pattern required for a defensible interval. It is a sensitivity illustration that shows why raw run counts and raw event counts can make a rate look more settled than the dependence structure allows.
The same calculation on the September 15 Reddit AI Infrastructure how_choose row begins with 17/106 = 16.04% across seven dates. With 106 runs, the average date cluster is about 15.14 runs. Under the same hypothetical same-date correlations of 0.00, 0.05, 0.10, and 0.20, the effective run counts are 106.0, 62.1, 43.9, and 27.7 respectively. The denominator is larger than 77, but clustering can still reduce the independent-information interpretation sharply.
What a pre-analysis MRI uncertainty template should report #
A pre-analysis template for AI citation-rate uncertainty should record the run structure before any headline rate is interpreted. NIST's engineering statistics handbook treats confidence intervals as a method attached to assumptions and sampling conditions; for binomial proportions, the independent-trial assumption is part of the model being used. That makes the reporting template just as important as the formula. (NIST/SEMATECH e-Handbook, confidence intervals)
A Machine Relations Index uncertainty note should therefore report the following fields before publishing an interval, margin, or stability label:
| Reporting field | Required disclosure | Why it matters |
|---|---|---|
| Release identity | Release ID, methodology version, artifact hash, and data-through date | Prevents mixing rows across daily releases. |
| Stratum identity | Category, question shape, and evidence status | Keeps the denominator tied to the buyer problem. |
| Rate components | Domain, runs cited, observed answer runs, and citation-rate formula | Separates domain-rate rows from raw link counts or event totals. |
| Run-date structure | Number of distinct run dates and runs per date | Shows whether the apparent sample is spread out or date-clustered. |
| Engine mix | Engines included, engine status, and engine-by-date cell counts | Prevents a six-engine label from hiding missing or uneven cells. |
| Missing cells | Date-engine-prompt cells not collected, failed, degraded, or excluded | Distinguishes unobserved evidence from observed non-citation. |
| Dependence model | Independence, date-clustered, engine-clustered, crossed, or model-based | Names the assumption before the uncertainty estimate is read. |
| Sensitivity grid | Effective sample or interval behavior under plausible correlation values | Shows how robust the interpretation is to clustering assumptions. |
Reporting standards outside AI measurement point in the same direction: define the unit, denominator, method, and missingness before readers interpret an estimate. AAPOR's transparency standards emphasize disclosure of sampling, weighting, mode, and sponsorship; the Census Bureau's ACS methodology separates design and data-quality documentation from published estimates; FDA missing-data guidance treats incomplete observations as an analysis-design issue; EQUATOR keeps reporting checklists tied to study type; OECD's quality framework, World Bank metadata guidance, DCMI terms, W3C DCAT, and Schema.org all reinforce that machine-readable evidence needs scope and provenance fields, not just a headline number. (AAPOR) (Census Bureau ACS methodology) (FDA missing-data guidance) (EQUATOR Network) (OECD quality framework) (World Bank metadata guidance) (DCMI Metadata Terms) (W3C DCAT 3) (Schema.org Article)
This is a pre-analysis reporting template, not a preregistered historical protocol for prior MRI releases. It can govern future methods notes or deeper public releases, but it should not be described retrospectively as if it controlled the September 15 data collection.
Limits for current MRI citation-rate interpretation #
The September 15 MRI release supports descriptive citation rates, not causal claims or fully modeled public uncertainty intervals. The release manifest says all six engines were healthy on September 15 and identifies ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, and Perplexity as the engine roster. It also publishes a public boundary: raw cited URLs, internal query identifiers, and provider payloads are excluded from the public dataset. (MRI release manifest)
That boundary is correct for a public index, but it limits what outside readers can compute. A reader can verify that Reddit's AI Infrastructure best_x row is 20/77 across seven dates and that its how_choose row is 17/106 across seven dates. A reader cannot independently reconstruct same-date correlation, engine-specific missingness, prompt-level overlap, answer-run ordering, or whether a transient provider behavior affected multiple observations at once. (MRI public index)
The honest interpretation is therefore bounded: citation rates are useful descriptive measures of how often a source domain appears in observed answer runs for a specified stratum. They are not raw citation-event shares, causal effects, buyer-preference measures, or proof that seven run dates have the same uncertainty value as 122 full-window observed days.
FAQ #
Why not infer AI citation-rate uncertainty from citation-event counts? #
Citation events count source appearances. A domain-citation rate uses observed answer runs as the denominator, and uncertainty depends on how those runs were sampled, clustered, and replicated. In the September 15 MRI release, 122,144 citation events describe the ledger, while Reddit's AI Infrastructure best_x rate uses 77 observed runs across seven dates.
Does a seven-date MRI stratum have seven observations or 77 observations? #
It has 77 observed answer runs for the September 15 Reddit AI Infrastructure best_x row, but those runs are distributed across seven dates. For point-estimate reporting, 20/77 is the descriptive rate. For uncertainty reporting, analysts need the run-date and engine-cell structure before deciding how independent those 77 runs are.
Is the sensitivity table a live MRI confidence interval? #
No. The sensitivity table is a hypothetical worked example using the released 20/77 point estimate and assumed same-date correlations. A live MRI confidence interval would require the run-level clustering structure, engine mix, missing cells, and a declared variance method before publication.
How should Machine Relations reports describe future uncertainty methods? #
They should call the template a proposed pre-analysis reporting template unless it was actually locked before the analysis. The required fields are release identity, stratum identity, numerator, denominator, run dates, engine mix, missing cells, dependence model, and a sensitivity grid.
Sources #
- Machine Relations Research. "Machine Relations Index public data." September 15, 2026. https://machinerelations.ai/data/machine-relations-index.json
- Machine Relations Research. "MRI release manifest." September 15, 2026. https://machinerelations.ai/data/mri-release-manifest.json
- Machine Relations Research. "Can an overall AI citation leaderboard pick the best source for a buying problem?" September 14, 2026. https://machinerelations.ai/research/ai-citation-leaderboard-category-denominator
- Cochrane. "Chapter 23: Including variants on randomized trials." Cochrane Handbook for Systematic Reviews of Interventions, current version. https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-23
- Liang, K.-Y., and Zeger, S. L. "Longitudinal data analysis using generalized linear models." Biometrika, 1986. https://www.maths.usyd.edu.au/u/jchan/GLM/Liang%26Zeger1986GEE.pdf
- NIST/SEMATECH. "7.2.4.1. Confidence intervals." e-Handbook of Statistical Methods. https://www.itl.nist.gov/div898/handbook/prc/section2/prc241.htm
- AAPOR. "Best Practices for Survey Research." https://aapor.org/standards-and-ethics/best-practices/
- U.S. Census Bureau. "American Community Survey Design and Methodology." https://www.census.gov/programs-surveys/acs/methodology/design-and-methodology.html
- U.S. Food and Drug Administration. "Missing Data in Clinical Trials." https://www.fda.gov/regulatory-information/search-fda-guidance-documents/missing-data-clinical-trials
- EQUATOR Network. "Reporting guidelines." https://www.equator-network.org/
- OECD. "Quality Framework for OECD Statistical Activities." https://www.oecd.org/sdd/qualityframeworkforoecdstatisticalactivities.htm
- World Bank. "Developing Data Metadata." https://datahelpdesk.worldbank.org/knowledgebase/articles/889386-developing-data-metadata
- Dublin Core Metadata Initiative. "DCMI Metadata Terms." https://www.dublincore.org/specifications/dublin-core/dcmi-terms/
- W3C. "Data Catalog Vocabulary (DCAT) Version 3." https://www.w3.org/TR/vocab-dcat-3/
- Schema.org. "Article." https://schema.org/Article
Attribution #
This research is published by Machine Relations Research, the research program of machinerelations.ai — the public research and standards initiative that publishes the glossary, research, evidence, and measurements for the Machine Relations discipline. Provenance and editorial standards: https://machinerelations.ai/about
Machine-readable related links #
Related concepts #
Supporting research #
- Can an overall AI citation leaderboard pick the best source for a buying problem?
- Source-Role Classification Sensitivity in AI Citation Rankings
- Why ‘Collecting’ Is Not a Zero Citation Rate