Between 2026-06-24 and 2026-09-22 we put the same buyer prompts to six AI engines every day. Five of them search the web before answering. One of them does not: it answers from what the model already holds.
On the prompts that did not contain our brand name, the engine that does not search mentioned us zero times. Not rarely. Zero, across 2,175 prompt-days.
Over the same 85 days, on the same prompts, on the same mornings, the five engines that do search moved us from 22.7 percent of the citation URLs they returned to 30.7 percent.
One number went up for three months. The other was never above zero to begin with.
What we measured #
The panel is our own instrument, run daily against a fixed registry of buyer prompts. It is described here in full because the result depends entirely on how the two layers are counted, and the two counts are not the same unit.
Six engines, split into two layers. Five are retrieval engines — Perplexity, ChatGPT, Gemini, Claude and Google AI Mode — each of which runs a search and returns an answer with citation URLs attached. Each of these vendors documents the retrieval step as an explicit tool: Anthropic's web search tool, OpenAI's web search guide, Perplexity's search control guide, and Google's documentation of AI features in Search and of AI Mode.
The sixth is the same Claude model with that tool switched off. No search, no search results block, no citations. It answers from its parameters. Anthropic's own framing when it shipped web search is the cleanest statement of the difference: retrieval exists because the model's own knowledge has a boundary. We call this sixth engine the parametric layer, and it is the control in this study.
Two denominators, deliberately different. On the retrieval layer we count citation URLs: every URL returned across every answered prompt is one slot, and a slot is ours if it points at a domain we own. On the parametric layer there are no URLs to count, so one answered prompt is one slot, and the slot is ours if the answer mentions us anywhere at all.
That asymmetry matters, and it runs against us on purpose. The parametric bar is the easier of the two. A passing mention clears it. On the retrieval layer nothing counts unless the engine returns a link to a page we own. We set the easier test on the layer that scored zero.
Scope. Errors are excluded rather than counted as absence: a prompt only enters a denominator on a day the engine actually answered it. Over the 85 measured days the parametric engine answered 2,768 prompt-days without error, so the zeros below are answers that did not mention us, not missing data. The registry ran 30 to 35 active prompts depending on the day.
The two layers, side by side #
| Window | Retrieval: owned share of returned citation URLs | Parametric: prompts mentioning us |
|---|---|---|
| 2026-06-24 to 2026-06-30 | 22.7% | 2 of 35 per day |
| 2026-08-08 to 2026-08-30 | 28.4% | 2 of 30 per day |
| 2026-09-08 to 2026-09-22 | 30.7% (2,495 of 8,151) | 3 of 482 prompt-days |
The retrieval column is the ordinary picture of a publishing programme working. We shipped assets into that window continuously, and the share of returned citation URLs that point at pages we own rose by roughly eight points and stayed there. On the last measured day the engines returned 552 citation URLs across the panel and 162 of them were ours.
The parametric column never participated. At its best it was two prompts out of thirty-something. Since 2026-09-08 it has been three mentions across 482 prompt-days.
Where the parametric memory actually lived #
The two prompts that ever produced a parametric mention are the two that already contain our name.
One asks who coined the term Machine Relations and names Jaxon Parrott inside the prompt. The other asks about a Machine Relations agency and names AuthorityTech inside the prompt. Both are recognition tests: the prompt supplies the entity, and the model confirms it.
Every prompt that does not supply the entity scored zero for 85 consecutive days. That includes five other prompts written around terms we coined and publish on — among them the plain definitional question about Machine Relations as a discipline, which named nothing of ours on any of the 85 days it ran.
This is the distinction interface research has made for decades and AI visibility measurement has not yet absorbed. Recognition is picking the right answer when it is placed in front of you; recall is producing it from nothing. A model that confirms your brand when the prompt hands it over is demonstrating recognition. A buyer asking a category question is testing recall.
On our panel the recognition rate at the parametric layer was 143 of 593 prompt-days, concentrated entirely in those two name-bearing prompts. The recall rate was zero of 2,175.
The regression we would have missed #
Recognition is not only distinct from recall. It is also less stable than it looks.
The prompt naming Jaxon Parrott resolved on 70 of the 85 measured days and last resolved on 2026-09-05. In the fifteen measured days since, it has not resolved once. The prompt naming AuthorityTech resolved on 73 of its 84 answered days and has resolved on only three of the fourteen days since 2026-09-08.
Nothing on the retrieval side moved with it. On 2026-09-07 — the first day the founder recognition prompt went dark — our retrieval share hit 34 percent, its highest reading in the entire 85-day series.
A brand's standing at these two layers is not one asset measured two ways. On these prompts, in this window, the layers moved in opposite directions on the same morning.
What this does to a share-of-voice number #
We publish the measurement contract for AI share of voice, and it defines the metric as mention breadth across a prompt set. This panel is that metric, computed honestly, and the result is that the number is almost entirely an artifact of which prompts are in the library.
Put our two name-bearing prompts in a thirty-prompt library and this brand reports a parametric share of voice near 6 percent. Remove them and the same brand, same day, same model, reports zero. Nothing about the brand changed. The prompt library changed.
That is not a reason to stop measuring. It is a reason to stop reporting one blended number. Three rules follow directly from this series, and each one is cheap:
Separate the layers before you average them. A score that pools retrieval citations with parametric mentions is adding a URL share to a mention rate. Report each on its own denominator, then roll up if you must.
Mark every prompt that contains your brand name. Those prompts measure recognition. They will flatter any brand that exists in training data at all, and they tell you nothing about whether a buyer asking a category question will meet you.
Watch the recall column for zeros, not for trends. Ours has no trend. It has a floor, and the floor is zero, and three months of publishing did not lift it off the floor even while retrieval standing rose eight points.
The commercial reading is the uncomfortable one. Retrieval standing is buyable with work and it compounds within days — that is what the first column shows, and it is the part of AI visibility that responds to what most teams already do. Parametric memory is not on the same clock. It is written when a model is trained, it is re-written when the model is replaced, and on the evidence here it can be lost in a week without anything on your side changing. This is why we treat independent corroboration across the wider source ecosystem as a different programme from citation work, and why the founder-level version of this finding was worth writing up separately.
It also sets the realistic expectation. With generative AI now used by a substantial share of US online adults and a majority of those users querying weekly, the answers buyers actually see are overwhelmingly retrieval-backed. That is the layer to compete on today. But every one of those answers is produced by a model that, asked the same question with the search tool off, does not know you exist.
Limits #
This is one brand, measured by an instrument we built and operate, against one parametric model. It is a single-subject longitudinal case, not a market measurement, and it does not establish that other models behave this way.
The absolute counts at the parametric layer are small: the largest daily footprint this brand ever had was two prompts. A move from two to zero is a move of two observations on any given day. The finding is not the size of the drop — it is that the floor held at zero across 2,175 prompt-days on every prompt that did not name us, and that the September loss on the two prompts that did name us has now persisted for more than two weeks.
The retrieval and parametric denominators are different units by construction, as described above. They are not directly comparable as percentages, and we do not compare them as such. What is comparable is direction, and the directions diverged.
The daily series behind every figure here is kept in the panel log that produces this page, and the panel continues to run. If the recall column ever leaves zero, that will be reported on this page.