Aggregate AI share of voice is a broken metric when it blends mentions, citations, recommendations, positions, engines, and query types into one unexplained number. An empirical test across five AI engines found that 77% of brands were cited by only one engine, and cross-platform citation data reported 11% domain overlap between ChatGPT and Perplexity. An 8,400-prompt study across four engines found cross-engine agreement on the top-cited brand reached 34% for head-term queries. Treating these engines as one measurement surface produces numbers that are directionally misleading unless the engine set, query set, counting rule, date window, weighting, and blended units are disclosed.
Last updated: June 23, 2026
The Divergence Problem: Six Engines, Six Citation Realities #
The core assumption behind aggregate share of voice is that AI engines behave like a single search channel with minor variation. The data refutes that at every level.
The Visionary Marketing study deployed 8,400 prompts across ChatGPT, Claude, Perplexity, and Gemini in February 2026 and reported that citation density varied by 44% between the most and least citation-dense engines:
| Engine | Brand citation rate | Citations per answer |
|---|---|---|
| Perplexity | 84.2% | ~21.9 |
| ChatGPT | 71.4% | ~7.9 |
| Gemini | 62.8% | ~17.0 |
| Claude | 58.4% | Prose-integrated |
Source: Visionary Marketing, AI Search Visibility Statistics 2026
Cross-engine agreement on the top-cited brand hit only 34% for head-term queries. For comparison queries ("best X for Y"), agreement dropped to 21.4%. A brand dominating one engine's citations can be invisible on another.
The VerityScore empirical test confirmed this pattern at the brand level. Using a single beauty-category prompt across five engines, 77% of cited brands appeared in only one engine's response. Only one brand — Clémence & Vivien — was cited unanimously by all five.
This is not just random noise. Each engine uses a different retrieval stack, different reranking logic, and different source preferences. Research on citation divergence shows that these architectural differences can produce systematically different citation patterns. Perplexity emphasizes reviews and expertise sources with heavy Reddit weighting. ChatGPT draws 48.7% of citations from third-party directories. Gemini favors brand-owned content at 52.2% — the opposite pattern.
Three Ways Aggregate Share of Voice Fails #
1. It Masks Engine-Level Concentration #
A brand with 25% aggregate SoV could hold 80% share on Perplexity and 0% everywhere else. An aggregate dashboard would call this "moderate visibility." An operator making budget decisions from that number would misallocate toward engines where the brand has no traction — or, worse, would neglect the single engine driving observable visibility.
The current Machine Relations Index exposes engine concentration as context rather than converting it into a global point score. In the public MRI v2 artifact, engine breadth records which monitored engines cited the root domain during the window. It helps an operator see whether citation behavior is distributed or concentrated, but it is not a weighted component in a universal authority score.
2. It Treats Mentions, Citations, and Recommendations as Equal #
Attrifast's analysis demonstrates the reporting problem: a definitional mention in a "what is X" answer counts differently from a source citation in a comparison answer, and both are different from an explicit product recommendation or a downstream business outcome.
The measurement repair is not to relabel every signal as AI SOV. Mention share, citation-source presence, recommendation language, position, traffic, and business lift are separate evidence units. A defensible visibility report labels each unit and discloses how, if at all, it is rolled into a broader AI Visibility Score.
3. It Cannot Detect Temporal Drift #
DigitalApplied's measurement framework documents 40–60% monthly shifts in cited domain sets within active categories. Ahrefs measured AI Overview citations from Google's top-10 results declining from 76% to 38% between July 2025 and March 2026.
A quarterly aggregate SoV snapshot captures a number that may have been true for only one of those three months. The current MRI public artifact records temporal context — days cited and days observed inside the artifact window — so a reader can distinguish a thin one-day observation from a repeated signal. That context describes observation coverage; it does not prove why the movement happened.
Per-Engine Measurement: The MRI v2 Contract #
The Machine Relations Index is built around per-engine and per-segment measurement because divergence makes unexplained aggregate tracking structurally unreliable. Its current methodology version is mri_score_v2.0.
MRI v2's unit is a root domain inside a source segment: a subject category plus a buyer-question type. For each segment, citation rate is calculated as cited answer runs divided by observed answer runs in that same segment. Thin segments remain collecting until the public artifact's evidence floor is met. As of the public artifact generated 2026-09-04, that floor is at least 10 observed runs across at least 7 distinct run dates.
The public MRI v2 artifact reports observed citation behavior. It does not publish a global weighted consensus score, a weighted authority score, a trust score, a recommendation score, a future-performance score, or a causal mechanism. Engine breadth, source role, position, temporal consistency, evidence state, confidence, and peer rank are context for interpreting the citation rate, not weighted point components.
| Field | What it records |
|---|---|
| Citation rate | Cited answer runs divided by observed answer runs in the same source segment |
| Evidence state | Published or collecting under the release evidence floor |
| Engine breadth | Monitored engines that cited the domain during the window |
| Source segment | Subject category plus buyer-question type |
| Position context | Average and best observed citation position |
| Temporal context | Days cited and days observed in the artifact window |
| Confidence | A, B, C, or collecting evidence-volume label |
As of the public artifact generated 2026-09-04, MRI v2 covered 20,658 cited root domains, 114,340 source events, 14,268 answer runs, 821 eligible prompts, and six engines, with an observation window from 2026-05-10 through 2026-09-04. The six engines in the artifact are ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, and Perplexity. The release manifest identifies the same methodology version and artifact window.
This per-engine granularity exposes patterns that aggregate measurement cannot:
- Distribution movement: Compatible-window reporting can show whether one engine's source share is rising or falling. Do not assert the cause of that movement unless a separate causal analysis exists.
- Engine-specific preferences: A current public row may expose a domain's citation rate, engine breadth, average or best position, source role, and segment context. Use those fields as observed citation behavior, not as proof that one engine will always cite a given role.
- Position context: If a domain appears higher on Perplexity than Gemini in a current row, treat that as a measurement prompt: inspect per-engine position context before averaging positions across engines.
What Operators Should Measure Instead #
The measurement gap is stark: only 14% of marketers currently track AI citations, while AI search visits grew 42.8% year over year from Q1 2025 to Q1 2026 (15.6B to 27.4B visits). The measurement infrastructure lags the channel growth by an order of magnitude.
A viable measurement stack for AI search visibility requires:
Per-engine citation tracking. Monitor each engine separately. The minimum viable engine set depends on the buyer surface being measured; MRI v2 currently observes ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, and Perplexity. Each has different retrieval architecture, source preferences, and update cadence.
Query-level granularity. Aggregate brand mention counts cannot distinguish between a definitional mention, a cited source, a comparison answer, and an explicit recommendation. Track which queries trigger mentions or citations and classify them by intent — informational, navigational, comparison, or transactional.
Temporal measurement cadence. Monthly snapshots can miss drift in citation sets. Weekly measurement gives operators a better chance to separate repeated movement from one-off model variance. The VerityScore methodology recommends minimum 3 runs per prompt to account for intra-engine variance of up to 15%.
Source role classification. Not all citations carry the same interpretation. The MRI classifies sources by public labels such as market database, analyst research, editorial publication, review platform, brand-owned, community forum, academic or standards source, and search or media platform. Source role helps explain observed patterns, but role alone does not determine which engine will cite a domain.
Position tracking, not just mention tracking. An 8,400-prompt study found that only 12% of AI citations go to brand-owned domains and reported a branded-search lift for brands receiving consistent mentions. Treat that as a study-specific, non-causal association. Position within a citation list may be useful context, but it should not be converted into user behavior, branded-search, traffic, or revenue without separate evidence.
Machine Relations Implication #
The measurement failure is not technical — the data exists. The failure is conceptual: treating architecturally distinct retrieval systems as one channel and counting mentions without decomposing them by engine, query, position, intent, and temporal stability.
AI Share of Voice is the competitive brand-mention breadth metric: qualifying brand mentions divided by all qualifying brand mentions in a fixed, declared panel. Share of Citation is a distinct scoped citation-source metric. The Machine Relations Index measures observed citation behavior by source segment. These three surfaces belong in the same measurement system, but they should not be collapsed into one undifferentiated number.
The evidence base for per-engine decomposition is no longer theoretical, but it is also not a single independent consensus claim. Practitioner and vendor analyses from VerityScore, Visionary Marketing, DigitalApplied, and Attrifast are directional evidence that adjacent parts of the AI-answer visibility problem require per-engine and per-unit disclosure. Pair that directional evidence with current public artifacts such as MRI v2, Microsoft Bing documentation, and Ahrefs' AI Overview research when making stronger operational claims.
Operators still reporting a single "AI share of voice" number to leadership should disclose the engine set, query set, counting rule, date window, weighting, and whether the number blends mentions, citations, recommendations, or positions. Without those disclosures, the number describes the reporting system more than the answer-engine surface.
FAQ #
What is AI share of voice? AI share of voice measures the percentage of qualifying brand mentions a company receives across AI-generated responses, relative to all qualifying brand mentions for the same declared category, engine, query, counting-rule, and time-window panel. It borrows from traditional media share of voice but applies the competitive mention-share idea to AI answer engines.
Why is aggregate AI share of voice unreliable? Because AI engines cite and mention dramatically different sources for the same queries. Empirical data shows 77% of brands were cited by only one engine, cross-platform overlap between ChatGPT and Perplexity was 11%, and cross-engine agreement on the top brand reached 34% for head-term queries. Aggregation is usable only when the engine set, query set, counting rule, date window, weighting, and blended units are disclosed.
How should brands measure AI search visibility instead?
Use per-engine tracking with query-level granularity, temporal cadence, source role classification, and position context. The current Machine Relations Index publishes mri_score_v2.0 citation rates by source segment, keeps thin segments in collecting, and uses confidence labels rather than a global weighted point score.
What is the difference between share of voice and share of citation? Share of voice counts qualifying brand mentions. Share of Citation is a separate citation-source metric under its own query, engine, matching, and time-window contract. Share of Citation complements AI SOV by showing citation-source presence or depth; it is not the same denominator and not the same outcome.
How often should AI citation data be measured? Weekly at minimum for active categories. Research shows 40–60% monthly shifts in cited domain sets within active categories, with up to 15% variance between identical runs on the same engine. Quarterly or monthly snapshots can miss most citation dynamics.