AI visibility scores are not comparable unless the platforms measure the same thing.
A score can represent the share of all responses that mention a brand, the share of responses that mention any brand, an impression-weighted competitive share, a composite of topic coverage and mention consistency, or a position-weighted brand score. Those are not interchangeable measurements. A number without its denominator is a decoration.
We reviewed public primary sources from 15 AI visibility and answer-engine tracking vendors on September 10, 2026. The sources included methodology pages, help centers, documentation, product FAQs, API references, and vendor-authored research. We did not use third-party roundups as evidence for what a vendor measures.
The result is a buyer's test: ask for the measurement contract before comparing the dashboard.
What the measurement contract must say #
A credible AI visibility measurement contract answers at least nine questions:
- What is the denominator for the headline score?
- How many times does each prompt run per engine and on what cadence?
- Which engines and model versions are included?
- Does the system record whether retrieval or web search happened?
- How does it expose sample size, variance, or confidence intervals?
- How are locale, personalization, and logged-in state controlled?
- Does it distinguish a brand mention from a linked citation?
- How are repeated mentions and citations deduplicated?
- Is the explanation a methodology document or product marketing copy?
If a vendor cannot answer these questions publicly, that does not prove its product is wrong. It does mean the score is not independently reconstructable from the public contract.
Three kinds of disagreement #
The audit found three different problems that are often collapsed into one.
Contradictory definitions #
A vendor publishes two definitions for what appears to be the same metric, with different denominators. That is a contract defect until the vendor explains which definition is current.
Complementary metrics #
A vendor publishes several metrics with different denominators, but names the objects clearly. Visibility, citation share, citation-total share, and position are different measurements. Different does not mean contradictory.
Undisclosed method #
A vendor publishes a score or feature but does not disclose enough of the denominator, sampling, uncertainty, or deduplication rule to reproduce it. The honest label is undisclosed, not inaccurate.
This distinction matters because a buyer can work with complementary metrics and ask a vendor to document an undisclosed one. A contradictory definition makes longitudinal comparison itself unstable.
What public vendor sources disclose #
Prominara: the clearest retrieval decomposition #
Prominara's public methodology states that every measured prompt runs multiple times per phrasing. It limits headline visibility to questions that do not mention the brand, distinguishes whether a brand was named from whether its site was linked, and records whether retrieval happened. Its citation framing decomposes the result as the probability that an engine searched multiplied by the probability of citation given search.
Prominara also describes intervals, stability readings, and control groups for impact testing. That is a stronger disclosure than a point score alone. It still leaves some exact run counts and locale controls undisclosed, but the objects being measured are comparatively clear.
Evertune: unusually explicit sampling #
Evertune's FAQ says it samples every prompt 100 times per model across more than 11 AI models. Its public material says the sampling is intended to reduce the noise of treating one response as a measurement and claims an approximate margin of error of one point overall and two points at topic level.
Its public definitions consistently describe Visibility Score as the percentage of AI responses that mention the brand. AI Brand Score adds position weighting. The sampling disclosure is unusually specific; the exact weights behind the composite score remain undisclosed.
Ahrefs: strong metric grain, conflicting share-of-voice language #
Ahrefs' AI visibility documentation defines a brand mention as a brand appearing at least once in a response and a citation as a page appearing at least once as a cited source. It also explains that a domain citation counts once when at least one page from that domain is cited, and that retrieved-but-not-cited pages can appear in a separate “Found in” measure.
The denominator problem is AI Share of Voice. The help material describes it as an impression share, with search volume weighting, while the product FAQ describes it as the percentage of AI responses in a topic set that mention or cite the brand versus competitors. Those can produce materially different values. The public pages do not say whether the product copy is shorthand for the weighted measure.
Scrunch AI: multiple denominators, clearly named #
Scrunch's citation-metrics documentation separates citation total share, citation share of voice, and citation rate. Citation total uses citation URLs as the denominator. Citation share of voice uses responses that cite any source. Citation rate uses all responses, including responses with no citation.
Its metrics guide and data-studio reference provide formulas based on response IDs and flags. The metric names can still be confused in a dashboard, but the public contract treats the differences as intentional rather than silently switching denominators.
Semrush: composite score, not a simple mention rate #
Semrush's AI Visibility Data documentation says its AI Visibility Score combines Topic Coverage and Mention Consistency. Its overview copy also uses simpler language about how often a brand appears, but the detailed description makes clear that the headline score is not merely brand-present responses divided by total responses.
Semrush says prompt responses are captured from real requests rather than through LLM APIs, and its Prompt Tracking feature can query target prompts by platform and location daily. Its AI Visibility Index methodology presents cross-platform results separately rather than hiding them inside one weighted number. Exact composite weights and a full uncertainty rule are not public in the material reviewed.
Peec AI: a reconstructable response denominator #
Peec's visibility documentation publishes a direct formula: responses mentioning the brand divided by total responses, multiplied by 100. Its API documentation describes the aggregation as the sum of visibility counts divided by the sum of visibility totals and gives a worked 5-of-10 example.
Peec says it runs prompts across platforms such as ChatGPT, Gemini, and Copilot daily. It also distinguishes sources—URLs accessed during response generation—from citations explicitly referenced in the answer. Exact runs per prompt and confidence treatment remain undisclosed.
Daydream: an explicit averaged leaderboard #
Daydream's leaderboard methodology says its ranking uses 12 questions per category, none naming a brand, and averages share across five AI tools rather than pooling all answers. A category page shows the run design as 12 questions, three phrasings, three runs, and five tools: 540 answers scored.
That is a useful disclosure because the aggregation rule changes the meaning of the final number. Daydream also separates the ranking from a citation trail built from additional competitor-named questions. The leaderboard is one dated run, not a rolling estimate, and it advises treating small differences as ties.
Retrieval is the missing variable #
Most public vendor documentation distinguishes mentions from citations. Far fewer say whether the answer engine actually searched or retrieved a source for a given response.
That omission matters. A cited page and a page merely present in the engine's retrieved source pool are different observations. A response generated from base-model knowledge is different from one generated after live search. A citation rate without retrieval exposure can make a change in search behavior look like a change in source quality.
Prominara explicitly records retrieval. Ahrefs exposes a “Found in” concept for pages retrieved but not cited. Evertune's public material distinguishes base-model and live-search behavior in its research and FAQ. Most of the other vendors reviewed expose cited URLs or source lists without a public per-run retrieval flag.
A buyer should therefore ask for the raw answer record, not only the score: prompt, engine, model or mode, timestamp, locale, retrieval state if available, answer text, cited URLs, and the deduplication rule applied to the metric.
Uncertainty is not optional #
AI answers vary by prompt phrasing, model, search state, time, location, and personalization. A single response is an observation, not a stable rate.
Prominara describes intervals and stability. Evertune describes repeated sampling and approximate margins of error. Scrunch and OtterlyAI discuss sample size and statistical significance. Several vendors warn that metrics are directional or that low prompt counts are unreliable. Many still publish a point score without a confidence interval or an explicit variance rule.
A methodology that reports 52% without saying whether it came from 10 responses, 100 responses, or 100 repeated runs per model is not giving the buyer enough information to compare 52% with another vendor's 52%.
The buyer's minimum request #
Before selecting or comparing an AI visibility platform, request one worked metric record and one worked aggregation:
- The exact prompt set and whether brand names were excluded.
- The engine, model, mode, locale, and logged-in state.
- The number of runs per prompt and collection cadence.
- The raw response count and the denominator used.
- Whether multiple mentions in one answer count once or many times.
- Whether multiple cited pages from one domain count once or many times.
- Whether retrieval/search happened and how it is recorded.
- The uncertainty treatment: interval, margin of error, stability band, or none.
- The exact formula for the headline score and every composite weight.
- The date and version of the methodology used for the report.
If the vendor cannot provide this, compare the platform's own trend over time under a frozen prompt set. Do not compare its headline score with another platform's headline score as though both were a common unit.
What this audit establishes—and what it does not #
This audit establishes that public AI visibility measurement contracts differ materially in denominator, sampling, retrieval disclosure, uncertainty treatment, and deduplication. It establishes that some vendors publish internally conflicting definitions and that others publish clearly different but complementary metrics.
It does not establish which vendor's score is most accurate in production. It does not show that more sampling produces better business decisions, that a disclosed methodology predicts citations, or that a particular platform improves traffic or revenue. Public documentation is evidence about what a vendor says it measures; it is not independent validation of the measurement's validity.
The practical rule is simple: buy the contract you can inspect, not the score you can screenshot.