A collecting label is not a zero citation rate. It means the published result is not yet eligible for interpretation under the dataset’s coverage rules. A measured zero is different: it requires an eligible stratum, observed runs, and no recorded citations in that denominator. A third state—not collectable—means the category cannot produce a valid rate in the current measurement design at all.
That distinction matters because a table can look complete while quietly turning missing evidence into negative evidence. Measurement guidance across official statistical and research methods sources treats the denominator and the missingness rule as part of the result, not as formatting details (NIST/SEMATECH e-Handbook, AAPOR standards, U.S. Census Bureau methodology, FDA missing-data guidance). The September 12, 2026 Machine Relations Index (MRI) release makes the boundary visible: 157 strata are classified as 79 published, 72 collecting, and 6 not collectable. The release covers 119 observed dates and 15,154 answer runs from May 10 through September 12, with a floor of 10 observed runs across seven dates. Those figures describe the release; they do not make every stratum equally supported. (MRI release manifest)
What the three states mean #
The first question is not “what is the rate?” It is “is a rate interpretable here?”
| State | What the dataset is saying | Can you report a citation rate? | Correct reading |
|---|---|---|---|
| Published | The stratum has passed the release’s publication floor and has an observed denominator | Yes, with its stated denominator and limits | A descriptive result for this eligible stratum |
| Collecting | The stratum is still accumulating enough eligible observations | No—not as a published rate yet | Evidence is incomplete, not zero |
| Not collectable | The current design cannot form a valid eligible denominator for the category | No | The category is outside this release’s rate calculation |
The labels are about evidence status, not performance tiers. Published does not mean certain, causal, representative of all engines, or predictive of future answers. It means the result cleared the release’s stated publication floor. Collecting does not mean the engines failed to cite the source. It means the release does not yet support that conclusion. Not collectable is not a poor score; it is a design boundary.
The public MRI index should therefore be read as a set of denominators with eligibility metadata, not as a leaderboard in which every blank is a zero. The release manifest records the methodology version and engine health, while the index carries the measured strata and their status. Keeping those layers together is what prevents a missing run from being silently counted as a non-citation. (MRI public index) (MRI research archive)
Published is eligible, not absolute #
A published stratum has enough observed evidence for the release to show a rate. It still has a bounded scope. The September 12 release has 15,154 answer runs across the whole May 10–September 12 window, but the floor is local: at least 10 observed runs across seven dates for a stratum to clear the publication boundary. A globally healthy engine roster does not guarantee balanced support for every category, question shape, date, or engine combination.
That is why a responsible result keeps at least four pieces attached to the number:
- The numerator: how many eligible runs contained the measured citation event.
- The denominator: how many eligible runs were actually observed for that stratum.
- The observation window: the dates represented in the release.
- The coverage caveat: which engines, prompts, or strata may still be unevenly represented.
A published zero, when one exists, would mean zero citations in that eligible denominator. It would not mean “the source is never cited,” “the engine rejects this source,” or “the source has no authority.” It would be a descriptive result for a defined run set. Even a clean zero cannot establish why the event did not occur. This is the same boundary used in observational and clinical reporting: an observed absence is not a causal explanation, and an incomplete record is not an observed absence (NIH/NLM guidance, CONSORT extension for harms, EQUATOR Network reporting resources). (MRI methodology reference)
Collecting is missingness, not a result #
A collecting stratum has a different logical shape. Its denominator is still being assembled or has not met the release’s eligibility floor. There may already be citations in the observed runs, but the release does not yet present that partial record as a published rate. Conversely, there may be no citations in the runs seen so far. Neither observation justifies writing “0%.”
The safe sentence is: “This stratum is collecting; no published rate is available in this release.”
The unsafe sentence is: “This stratum has a zero citation rate.”
Those sentences differ by only a few words, but they make opposite claims about the evidence. The first preserves the missingness state. The second asserts a measured outcome and implies a denominator that the release has not granted.
This is also why pooling collecting strata into a larger average is not a harmless convenience. A combined number can hide which rows are eligible, which are still accumulating, and which cannot be measured. The question-shape pooling note proposed on September 11 addresses comparability and aggregation; this note addresses the prior decision: whether a row is a result at all. Do not substitute one method question for the other. (MRI source-layer baseline)
Not collectable is a design boundary #
A not-collectable category cannot yield a valid rate under the current collection and classification design. That may happen because the category lacks a stable denominator, cannot be mapped to the required observation unit, or does not produce the kind of event the release is designed to measure. The important point is that the category is not waiting for “a few more runs” in the same sense as a collecting row.
The correct treatment is to retain the label and explain the boundary. Do not rank it below published strata. Do not impute zero. Do not turn it into a recommendation about an engine, a source, or a buyer. A data system can be healthy overall while still leaving particular categories unmeasurable; aggregate engine health is not proof of balanced support for every stratum. (MRI release manifest)
A hypothetical zero, kept separate from the release #
Consider a hypothetical stratum called example-question-set, with 20 eligible observed runs across seven dates. Suppose the measured source appears in 0 of those 20 runs, and the stratum is marked published because it satisfies the release’s eligibility rules.
The accurate statement would be:
In this hypothetical eligible stratum, the source was cited in 0 of 20 observed runs, a descriptive incidence of 0% for that defined window.
That is a measured zero within the example’s eligible denominator. It is not a statement that the source will never be cited, that an engine caused the absence, or that the source is generally unfit for retrieval. It also is not one of the 72 collecting strata or 6 not-collectable categories in the September 12 release. The example is deliberately hypothetical because the released brief does not supply a measured zero to borrow.
A practical decision tree #
Use this order before writing any rate:
- Is the category marked not collectable? Report the design limitation and stop. There is no valid rate in this release.
- Is it marked collecting? Report that evidence is incomplete and stop. Do not calculate or imply zero.
- Is it marked published? Read the local numerator, denominator, date coverage, and caveats.
- Is the numerator zero? Call it a measured zero only if the denominator is eligible and the release actually reports that zero.
- Are you comparing rows? Check that their question shape, engines, dates, and denominators are comparable before pooling or ranking them.
This sequence keeps status, measurement, and interpretation separate. It also makes the table auditable: a reader can see whether a non-result reflects incomplete collection, a true zero in an eligible window, or a category the design cannot measure.
What the release can—and cannot—support #
The September 12 release supports descriptive statements about the 157 classified strata and their release status. It supports the observation that 79 are published, 72 are collecting, and 6 are not collectable. It supports the stated whole-window run and date totals. It does not, by itself, support claims about buyer behavior, causal engine or source effects, or future citation performance. The latest Search Console context is also lagged: the August 11–September 8 export was generated September 11, with the ordinary three-day delay. Those impressions are context for reader demand, not a current-market sample.
The discipline is simple: a missing row is not a zero row. Publish a number only when the dataset publishes an eligible denominator. Preserve collecting when the evidence is unfinished. Preserve not collectable when the measurement design stops there. That is not a softer way to report results; it is the difference between an interpretable measurement and a fabricated conclusion. For further primary references on distinguishing data availability from measured outcomes, see the OECD data quality framework, the World Bank metadata guidance, and the ISO data quality vocabulary.
FAQ #
Does collecting mean the source has no citations? #
No. Collecting means the release does not yet publish an interpretable rate for that stratum. Partial observations may exist, but they should not be presented as a completed denominator or a zero.
Can a published stratum have a zero citation rate? #
Yes, if the stratum is eligible, its denominator is defined, and the released result records zero citations in that denominator. That is a descriptive zero for the stated window—not proof of permanent absence or causation.
What should I do with a not-collectable category? #
Keep the category visible, label the design limitation, and exclude it from rate rankings and averages. Do not convert it to zero or treat it as a low-performing published stratum.
Does the publication floor prove statistical certainty? #
No. The floor is a release eligibility rule. It makes a result publishable under the stated methodology; it does not turn a bounded observational dataset into statistical certainty.
Last updated: September 12, 2026.