# Three Daily Releases, One Unchanged Leaderboard: What Our Index Stops Observing After a Segment Publishes

We compared two Machine Relations Index releases three days apart, segment by segment. All 84 comparable published segments were numerically identical — same observed runs, same run dates, same rank-1 citation rate. Here is why, and what it means for anyone buying a daily AI visibility dashboard.

Canonical URL: https://machinerelations.ai/research/ai-citation-leaderboard-movement-between-releases-2026
Published: 2026-09-21
Research type: Index Analysis
Tags: measurement, citations, ai-search, methodology, machine-relations-index

## Source Body

The Machine Relations Index regenerates every day. Between the release generated on September 18, 2026 and the one generated on September 21, it added 198 cited domains, 2,527 source events and 364 observed answer runs.

Across those same three days, not one of the 84 segment leaderboards published in the earlier release changed by a single observation.

Not "changed a little." The same number of observed runs, over the same number of distinct dates, with the same citation rate at rank 1. Eighty-four out of eighty-four. We went looking for the movers and found that the category's most basic assumption about our own instrument — that a daily release means a daily measurement — is not true of the segments anyone actually reads.

This is the arithmetic, the reason, and what we are doing about it.

## What we compared

A **stratum** in the Index is one subject category paired with one question shape — `cybersecurity` asked as `how_choose`, `consumer-finance` asked as `top_list`, and so on across 25 taxonomy nodes and six question shapes. Every segment leaderboard we publish, and every buyer shortlist read off one, is a stratum.

We took two releases:

| | September 18 | September 21 |
|---|---|---|
| Release id | `mri_score_v2.0+2026-09-18+8fa38e54dd0a` | `mri_score_v2.0+2026-09-21+d73e239e7c5b` |
| Window | 2026-05-10 → 2026-09-18 | 2026-05-10 → 2026-09-21 |
| Days observed | 125 | 128 |
| Cited domains | 22,179 | 22,377 |
| Source events | 124,397 | 126,924 |
| Observed answer runs | 15,782 | 16,146 |
| Cited-domain observations | 33,756 | 34,184 |
| Published strata | 85 | 91 |
| Collecting strata | 66 | 60 |

The Index plainly grew. Three more days of observation, 198 more domains, 364 more answer runs.

Eighty-five strata were published in the earlier release. Eighty-four of them carry the same key in both. (The eighty-fifth was filed under a placeholder question-shape key that has since resolved to `deep-tech-hardware | is_x_worth`; it is excluded from the comparison rather than counted as a loss.)

For each of those 84 we compared three things that cannot move unless the segment was observed again: the number of observed runs, the number of distinct run dates, and the citation rate of the domain at rank 1.

## The result

All 84 were identical on all three.

Summed across the 84 shared segments, observed runs total **14,614 in both releases**. Not 14,614 and 14,620. The same number.

To be concrete about what that looks like on a single segment: in `ai-infrastructure | best_x`, reddit.com sits at a 0.2597 citation rate on 20 cited runs out of 77 observed, across 7 distinct dates. That is its line in the September 18 release. It is character-for-character its line in the September 21 release.

## Why: the seven-date wall

The Index publishes a stratum only after it clears an **evidence floor** — at least 10 observed runs across at least 7 distinct dates. Below that line a stratum is marked `collecting` rather than scored, so thin signal is never shown as settled. That floor is a good rule and we have [published research defending it](/research/ai-citation-evidence-floor-scoreable-domains-2026).

Here is the run-date distribution across all 91 currently published strata:

| Distinct run dates | Published strata |
|---|---|
| 7 | 84 |
| 53 | 5 |
| 54 | 2 |

**Eighty-four of ninety-one published segments sit at exactly seven run dates** — precisely the floor's minimum, not one date more.

And the seven exceptions are not a random sample. Every one of them is a `news_topic` bucket:

| Stratum | Run dates | Observed runs |
|---|---|---|
| legacy-unmapped \| news_topic | 54 | 2,138 |
| cybersecurity \| news_topic | 54 | 623 |
| fintech \| news_topic | 53 | 616 |
| healthcare-services \| news_topic | 53 | 613 |
| hr-talent \| news_topic | 53 | 608 |
| martech-advertising \| news_topic | 53 | 602 |
| enterprise-software \| news_topic | 53 | 596 |

So the legacy news buckets are under continuous observation. Every buyer-decision segment — the 24 subject categories crossed with the six question shapes, which is the entire evidence base for a shortlist, a leaderboard or a "who does AI cite when someone is choosing a vendor" answer — is a **one-time seven-date snapshot that stops accumulating the moment it publishes**.

The published methodology describes the evidence floor as a threshold for *publication*. It does not describe any mechanism that stops *observation* once a stratum crosses it. We are reporting what our own public data shows, not a behaviour our stated method predicts.

## Where the new observations went

Into strata that had not published yet.

Six genuinely new segments crossed the floor in this window, and they arrived exactly where you would now expect — at seven run dates each:

| Newly published stratum | Run dates | Observed runs |
|---|---|---|
| healthcare-services \| problem_first | 7 | 108 |
| healthcare-services \| best_x | 7 | 107 |
| healthcare-services \| how_choose | 7 | 107 |
| healthcare-services \| is_x_worth | 7 | 107 |
| healthcare-services \| top_list | 7 | 102 |
| hr-talent \| how_choose | 7 | 101 |

That is the full mechanism. A stratum accumulates while it is `collecting`, crosses the floor, publishes — and then goes quiet. Collecting strata fell from 66 to 60 over the same three days. The growth is real and it is entirely in front of the floor, never behind it.

The head of the Index behaves the same way. Of the top 25 domains by publisher rank, 24 are unchanged; deloitte.com left the top 25 and guideflow.com entered. Citation rates drift very slightly *down* across the head, because the denominator grew and the new runs were largely cited by other domains. Reddit was cited in 9 of the 364 new runs — a 2.5% hit rate against its 12.52% standing rate.

| Domain | Rate 09-18 | Rate 09-21 | Cited runs 09-18 | Cited runs 09-21 |
|---|---|---|---|---|
| reddit.com | 0.1275 | 0.1252 | 2,012 | 2,021 |
| youtube.com | 0.0907 | 0.0889 | 1,431 | 1,436 |
| linkedin.com | 0.0681 | 0.0679 | 1,075 | 1,096 |
| medium.com | 0.0511 | 0.0500 | 807 | 808 |
| forbes.com | 0.0411 | 0.0412 | 649 | 666 |

## The movement that looked real was a tiebreak

Three segments did show a different domain at rank 1, and fifteen showed a different top three. We checked every one of them before writing this, and none is movement.

In each case the two domains that swapped carry **identical citation rates and identical cited-run counts**:

| Stratum | Rank 1 | Rank 2 | Both at |
|---|---|---|---|
| consumer-products \| top_list | chewy.com | dogfoodadvisor.com | 0.2366, 31 cited runs |
| emergent-prosumer \| best_x | reddit.com | techradar.com | 0.2100, 21 cited runs |
| emergent-prosumer \| top_list | digitalcameraworld.com | reddit.com | 0.1980, 20 of 101 runs |
| family-software \| how_choose | apple.com | reddit.com | 0.1731, 18 of 104 runs |

That prompted a wider count. Across all 91 published strata:

- In **4**, rank 1 and rank 2 are numerically indistinguishable.
- In **11**, a tie sits somewhere inside the top three.
- In **86 of 91**, a tie sits somewhere inside the top ten.

The release resolves every tie into distinct ranks — no stratum shows two domains sharing rank 1. Which means the ordinal is doing work the measurement cannot support. On four published leaderboards, a brand reading "we are #1 and they are #2" is reading a tiebreak. On 86 of 91, at least one adjacent pair in the top ten is separated by nothing the instrument can see.

This is the more actionable of the two findings. A rank gap of one position is not evidence of anything unless you also know the citation counts behind it, which is why every segment page in the Index prints the rate and the cited-run count next to the rank.

## Why this is not a contradiction of the volatility research

The obvious objection: a large body of measurement says AI citations churn constantly. It does, and we are not disputing it.

[GetMentions ran 2,398 queries across four engines once a day for seven consecutive days in June 2026](https://www.getmentions.ai/blog/ai-citation-volatility-study) — 67,144 answers and 530,875 citations — and found day-over-day source churn from 44.4% on Perplexity to 88.3% on Gemini. Their one-line conclusion is that AI visibility "is not a rank you hold. It is a probability." [SISTRIX analysed 82,619 prompts over 17 weeks](https://www.sistrix.com/blog/ai-citation-drift-how-stable-are-sources-in-ai-search-results/) across three platforms and six countries: Google AI Mode replaces 56% of its cited sources weekly and ChatGPT 74%, with drift rates holding at 54–59% across every country and showing no sign of settling. [Parse measured repeat answers to the same prompt](https://www.parse.gl/research/ai-citation-volatility-by-industry) and found ChatGPT shared only 21.2% of cited domains between two runs of an identical prompt, against 31.5% for Google AI Overviews. [Trakkr ran the same brand queries daily for ten months across 10,000 brands](https://trakkr.ai/trakkr-research/citation-decay) and found 70.2% of cited URLs appeared exactly once, with 29 days from peak to half. [FogTrail's continuous monitoring](https://fogtrail.ai/blog/how-often-ai-search-engines-update-citations) puts the refresh cycle at roughly 48 hours, with turnover per cycle ranging from under 5% on ChatGPT to over 40% on Perplexity. The debate itself was opened by [Profound's drift analysis](https://www.tryprofound.com/blog/ai-search-volatility) of roughly 80,000 prompts per platform, comparing the same open-ended queries a month apart. A Writesonic study of 23 million cited sources from April to June 2026, [reported by Foundation](https://foundationinc.co/lab/vol-298), put the typical citation lifespan at 11 to 15 days with 44% of cited pages appearing exactly once. Closest to the ranking question this page is about, [LumenGEO's maintained statistics roundup](https://lumengeo.co/blog/ai-search-statistics-2026) puts top-ten stability across identical re-asks at roughly 89% — about one source in nine changing inside the part of the answer everyone watches.

And at the domain layer, [BrightEdge found the opposite kind of number](https://www.brightedge.com/resources/weekly-ai-search-insights/ai-search-citations-week-to-week-changes): 96.8% of cited domains saw zero week-over-week change, rising to 99.4% for brands in the #1 or #2 mention position. We have [synthesised](/research/ai-citation-stability-week-to-week-evidence-2026) that [apparent paradox](/research/ai-search-citation-volatility-weekly-stability-2026) before: domains are sticky, individual URLs churn beneath them.

None of that is what we measured here. Every one of those studies re-asked the questions. Our zero is not a finding about how AI engines behave over three days — it is a finding about what our own release re-observed over three days, which is nothing, in any segment that had already published.

That distinction is the whole point, and it cuts at the entire category. **A stability number is meaningless unless you know the instrument re-measured.** "96.8% unchanged" and "84 of 84 unchanged" look like the same kind of claim and are not: the first is a measurement of the world, the second is a measurement of a measurement. Any AI visibility dashboard quoting you a week-over-week stability figure owes you the answer to one question before the number means anything — how many times did you re-ask, and on which dates?

We could not answer that question about our own segments until we ran this comparison. That is the uncomfortable part, and it is why this page exists.

## What this does and does not change

**It does not change any published citation rate.** Every rate, cited-run count, date count and confidence grade in both releases is an accurate report of the runs that were actually observed. Nothing here is a correction; no number is being withdrawn. A seven-date sample that cleared the evidence floor is exactly what we said it was.

**It does change what a release date means.** A segment leaderboard carrying today's release id is not a measurement taken today. It is a measurement taken across seven dates, some time before the stratum published, and reprinted under every release id since. We have not been claiming otherwise in words, but publishing under a daily release id implies a recency that 84 of our 91 published segments do not have.

**It changes what we will ship.** We had a weekly "index report" queued to run on release-over-release movers, and a per-engine change log waiting for a first real movement to record. Both were waiting on movement that the instrument, as currently run, cannot produce in a published segment. Neither will ship as designed. The honest version of a change log is one that records when a segment is *re-observed*, not when a release id increments.

The fix belongs upstream of editorial: published strata need to keep accumulating observations the way the news buckets do. Until they do, we will state the observation window on segment pages rather than let the release date stand in for it.

## Check it yourself

Everything above is derived from the public release artifact, which is the same file we read:

- The machine-readable release: `https://machinerelations.ai/data/machine-relations-index.json
- Any segment leaderboard, for example [`/index/categories/ai-infrastructure/best_x`](https://machinerelations.ai/index/categories/ai-infrastructure/best_x), which prints the rate and cited-run count beside every rank
- Any domain profile, for example [`/index/domains/reddit.com`](https://machinerelations.ai/index/domains/reddit.com), which lists every stratum the domain appears in with its rank and rate

Each published stratum carries `runs_observed` and `run_dates` in `mri_score_v2.strata[]`. Pull two releases a few days apart and compare those two fields on any segment that was already published in the earlier one. You will get what we got.

Corrections to this analysis, or to anything else we publish, go in the [Index correction record](/research/ai-citation-index-correction-record-2026).

## Related reading

- [The evidence floor and the scoreable population of the Index](/research/ai-citation-evidence-floor-scoreable-domains-2026) — why only a fraction of cited domains carry a confidence grade
- [How stable are AI search citations week to week](/research/ai-citation-stability-week-to-week-evidence-2026) — the two-tier domain/URL structure
- [Citation stability by source role](/research/citation-stability-source-role-patterns-ai-search-2026) — which source types hold citations longest
- [Why AI visibility scores differ between tools](https://paralabs.ai/blog/why-ai-visibility-scores-differ-between-tools) — how much each methodology choice moves a score
- [AI citation rate benchmarks by industry and brand size](https://authoritytech.io/curated/ai-citation-rate-benchmarks-industry-brand-size-2026) — the published distribution to place your own number in
- [HR software leaderboard rank-distance](https://paralabs.ai/blog/hr-software-ai-leaderboard-rank-distance-2026) — what the gap between rank 1 and rank 10 is actually worth
- [Is one AI shortlist enough market proof](https://jaxonparrott.com/blog/founder-decision-ledger-hr-ai-shortlist-market-proof-2026) — the founder-side version of the same caution
- [Why press coverage does not reach the AI buying shortlist](https://christianlehman.com/blog/news-pitch-list-vs-ai-buying-answer-2026) — the news/buyer split these strata make visible

## Attribution

This research is published by Machine Relations Research, the research program of machinerelations.ai — the public research and standards initiative that publishes the glossary, research, evidence, and measurements for the Machine Relations discipline. Provenance and editorial standards: https://machinerelations.ai/about

## Machine-readable related links

### Related concepts

- [Machine Relations Index (MRI)](https://machinerelations.ai/glossary/machine-relations-index)
- [MRI Score](https://machinerelations.ai/glossary/mri-score)
- [Machine Relations (MR)](https://machinerelations.ai/glossary/machine-relations)
- [AI Visibility](https://machinerelations.ai/glossary/ai-visibility)

### Supporting research

- [Only 508 of 22,377 AI-Cited Domains Are Cited Often Enough to Score](https://machinerelations.ai/research/ai-citation-evidence-floor-scoreable-domains-2026)
- [When an AI Citation Index Is Wrong: One Misclassified Field, 57 Corrected Pages](https://machinerelations.ai/research/ai-citation-index-correction-record-2026)
- [Ranked Without a Grade: Half of the AI Citation Index's Top-Ten Positions Belong to Ungraded Domains](https://machinerelations.ai/research/ranked-without-a-grade-ai-citation-index-domain-lookup-2026)
- [How Stable Are AI Search Citations Week to Week? Evidence From Three Measurement Systems](https://machinerelations.ai/research/ai-citation-stability-week-to-week-evidence-2026)

### Framework context

- [Machine Relations Index](https://machinerelations.ai/index)
- [Machine Relations Stack](https://machinerelations.ai/stack)
- [Evidence Base](https://machinerelations.ai/evidence)
