Research

AI Visibility Measurement: How B2B Brands Track What AI Engines Actually Cite

A reference guide to measuring AI visibility for B2B brands — comparing prompt-based monitoring tools, source-level citation rate measurement, and the metrics that connect AI engine behavior to pipeline.

Published Machine Relations Research
Reference

AI visibility measurement tracks how often, how prominently, and through which sources AI engines surface a brand when buyers ask questions about its category. The practice splits into two distinct disciplines: prompt-based brand monitoring (does the AI mention you?) and source-level citation measurement (does the AI cite your content as evidence?). Both matter. Most tools handle the first. Few handle the second. This guide covers the measurement stack, compares leading tools, and explains what citation rate data reveals that brand-mention tracking cannot.

Why Traditional SEO Metrics Miss AI Visibility #

Traditional search measurement — keyword rankings, organic traffic, click-through rates — was built for a page of ten blue links. AI engines work differently. ChatGPT, Perplexity, Google AI Overviews, Claude, Gemini, and Microsoft Copilot synthesize answers from multiple sources, cite some of them, mention brands by name without linking, and vary their responses across sessions.

A brand can rank #1 on Google for a query and never appear in the AI-generated answer for that same query. Nobori's 2026 analysis found that median citation inclusion rates in AI Overviews sit around 3% — meaning most brands that rank organically are not cited when AI answers the same question. The Otterly AI Citations Report 2026 found 73% of websites block AI crawlers through robots.txt or CDN rules, cutting off the supply line before measurement begins.

These gaps mean brands need a distinct measurement layer for AI. Porting SEO dashboards into AI reporting produces blind spots.

Two Measurement Disciplines: Brand Monitoring vs. Source Citation #

AI visibility measurement splits along a line that most tool comparisons blur.

Prompt-based brand monitoring answers: "When someone asks an AI engine a question in my category, does it mention my brand?" Tools in this category run structured prompts against AI engines, record whether the brand appears, and compute a visibility score or share-of-voice percentage. Semrush AI Visibility Toolkit, Otterly.ai, Profound, Peec AI, AthenaHQ, and Cognizo operate here. The output is directional: up, down, or stable for a given set of prompts.

Source-level citation measurement answers a different question: "How often do AI engines cite a specific source domain as evidence, across which queries, engines, and verticals?" This measures what the engines do with sources — not what they say about brands. The Machine Relations Index (MRI) operates here, measuring citation rates across six engines for over 6,000 domains, publishing rates only after a segment clears an evidence floor of at least 10 observations across at least 7 distinct run dates.

The distinction matters because brand mentions and source citations are different events. A brand can be mentioned without any source being cited. A source can be cited without the brand being named. Measuring one without the other leaves half the picture unmapped.

Core Metrics for AI Visibility Tracking #

The metrics that matter depend on which discipline you are measuring. Here are the metrics used across both, with what each actually tells you.

Metric What It Measures Typical Source Limitation
Visibility Score % of tracked prompts where a brand appears Cognizo, Otterly, most tools Depends entirely on prompt selection; different prompt sets produce different scores
Share of Voice Brand's mention share relative to competitors Semrush, AirOps, Peec AI Requires defining the competitive set; sensitive to which competitors are included
Citation Rate % of observed AI answer runs citing a specific domain Machine Relations Index Requires large observation volume and cross-date consistency for stability
Prominence Score Weighted position of brand within the AI answer Yolando Scoring weights are arbitrary across vendors; no industry standard
Sentiment Alignment Whether AI describes the brand accurately Gracker, Scrunch Binary or coarse scales; hard to act on without content-level attribution
AI-Sourced Pipeline Revenue attributable to AI-referred traffic GA4 + UTM tagging ChatGPT appends utm_source=chatgpt.com; Perplexity and Claude do not yet

Yolando's metrics framework recommends running 60–100 prompt executions per query for statistical significance. Maximus Labs specifies minimum 30 sampling runs per query per platform with 95% confidence intervals. These thresholds exist because AI responses vary between sessions — a single prompt run is an anecdote, not data.

AI Visibility Tools Compared: 2026 Vendor Landscape #

The tool market has expanded rapidly. Here is a comparison of leading platforms by measurement capability, engine coverage, and pricing, verified as of mid-2026.

Tool Engines Tracked Primary Metric Entry Pricing Differentiator
Profound 11 platforms incl. ChatGPT, Claude, DeepSeek, Amazon Rufus Citation share, product visibility $99/mo (ChatGPT only); enterprise custom $96M Series C at $1B valuation (Feb 2026); agent analytics
Semrush AI Toolkit ChatGPT, Perplexity, Gemini, AIO Share of Voice weighted by prompt volume $99/mo add-on (effective floor ~$239 with Pro) 25-prompt cap; integrates with existing SEO stack
Otterly.ai ChatGPT, Perplexity, AIO, Gemini Brand Visibility Index $29/mo Lite (15 prompts) – $489/mo Gartner Cool Vendor 2025; link citation detection
Peec AI ChatGPT, Perplexity, Google AI Real-time visibility alerts €89 – €499+/mo $29M raised, $4M+ ARR in 10 months; 115+ languages
AthenaHQ ChatGPT, Gemini, Claude, Perplexity, Copilot, AIO QVEM model; ACE Citation Engine $270 – $2,000+/mo Y Combinator backed; case study: Grüns 2.0% → 12.6%
Cognizo 10 engines incl. Grok, Meta AI, DeepSeek Visibility Score, owned vs earned citations Custom Separates owned citations (your domain) from earned (third-party)
TurboAudit ChatGPT, Perplexity, Gemini Audit + monitoring combo Free / $39.99/mo 250+ AI checks; low entry price
HubSpot AEO Grader ChatGPT (GPT-5.4 mini), Perplexity, Gemini AEO Score (10-point SoV component) Free 100 test queries; single-shot baseline, not continuous
Machine Relations Index ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, Perplexity Citation rate per domain per segment Published research (public data) 6,020+ domains measured; evidence floor before rates publish; confidence tiers

Most tools operate on the same basic model: define a prompt set, run it against AI engines on a schedule, record brand mentions, compute scores. The differences lie in engine coverage breadth, statistical rigor of sampling, and whether the tool distinguishes citation (source linked) from mention (brand named without attribution).

What Citation Rate Measurement Reveals That Brand Monitoring Cannot #

Brand monitoring answers whether AI knows your name. Citation rate measurement answers whether AI trusts your content enough to use it as evidence.

The Machine Relations Index v2 measures source-segment citation rates — how often AI answer engines cite each source domain within a specific subject category paired with a buyer question type. Rates publish only after the segment clears an evidence floor: at least 10 observations across at least 7 distinct run dates. Below the floor, a domain is "collecting," never scored. Above the floor, each domain carries a confidence grade — A, B, or C — reflecting how much evidence stands behind its rate.

This approach reveals patterns that prompt-based monitoring misses:

Cross-engine disagreement. Different AI engines cite different sources for the same question type. A domain may hold a high citation rate on Gemini but appear rarely on Perplexity. Brand monitoring would show "visible on both" if the brand is mentioned in either. Citation rate measurement shows which engine actually treats the domain as a source — and whether that treatment is stable across measurement dates.

Source-type stratification. The MRI classifies sources by role: editorial publication, market database, analyst research, government source, and others. Citation rates vary systematically by source type. A market database like G2 or Crunchbase earns citations through a different mechanism than an analyst firm like McKinsey or Gartner. Understanding which source type your domain belongs to — and how that type is treated — shapes what "improvement" even means.

Temporal stability. A one-time prompt check produces a snapshot. Citation rates measured across 7+ distinct dates and 10+ observations reveal whether a domain's citation authority is stable, rising, or eroding. A brand might appear in 40% of prompts today and 5% next week if the underlying source lost citation authority. Only longitudinal source measurement catches the structural shift.

Building a Measurement Framework: What to Track and When #

A practical AI visibility measurement framework combines both disciplines. Here is a phased approach based on what the data supports.

Phase 1: Baseline (Week 1–2). Start with a free single-shot audit. HubSpot AEO Grader runs 100 test queries across ChatGPT, Perplexity, and Gemini at no cost. TurboAudit offers a free tier with 250+ checks. The goal is not precision — it is a directional baseline: does AI know your brand, and for which questions?

Phase 2: Continuous brand monitoring (Month 1–3). Select 30–50 prompts covering your core buyer queries. Mix unbranded category queries ("best [category] tools 2026"), branded comparison queries ("[your brand] vs [competitor]"), and problem-solution queries ("how to solve [buyer problem]"). Run these weekly across ChatGPT, Perplexity, Gemini, and Google AI Overviews at minimum. Track visibility score, share of voice, and which competitors appear alongside you. Cognizo recommends tracking owned citations (links to your domain) and earned citations (third-party mentions) separately — this distinction matters for diagnosing what drives visibility changes.

Phase 3: Source-level citation analysis (Month 3+). Layer in source-domain measurement. Check which of your URLs AI engines actually cite using server log analysis (filter for ChatGPT-User, GPTBot, PerplexityBot, ClaudeBot, OAI-SearchBot user agents), or use the MRI public dataset to see how your domain's citation rate compares within its source type. This reveals whether your content is structurally citable — not just whether your brand is known.

Phase 4: Pipeline attribution (Ongoing). Connect AI visibility to revenue. ChatGPT appends utm_source=chatgpt.com to outbound links. Set up GA4 segments for AI-referred traffic. Gracker's cybersecurity report recommends tracking "AI-sourced pipeline" as the executive metric. Correlate branded search volume lifts with AI mention frequency to capture zero-click influence where no referral link exists.

Crawler Access: The Measurement Prerequisite Most Brands Fail #

Before any measurement framework matters, AI engines need to be able to reach your content. The Otterly 2026 report found 73% of websites block AI crawlers. This is the single highest-leverage fix: if GPTBot, PerplexityBot, ClaudeBot, and OAI-SearchBot are blocked in your robots.txt or at the CDN level, you have eliminated your content from the citation supply before measurement begins.

Check your robots.txt for these user agents. Check your CDN (Cloudflare, Akamai, Fastly) for bot-blocking rules that may catch AI crawlers as a side effect of general bot protection. A brand scoring 0% on visibility may not have a content problem — it may have an access problem.

Owned vs. Earned Citations: Measuring Both Sides #

Cognizo's framework introduced a distinction that clarifies measurement: owned citations link directly to your domain; earned citations come from third-party sources that mention your brand.

The VisibleIQ 2026 study found approximately 79% of citations on Perplexity, Gemini, and Claude come from third-party domains — not from the brand's own site. This means a brand's AI visibility depends heavily on what others publish about it. Measuring only owned citations misses the majority of the citation supply.

For B2B brands, this has direct implications. Press coverage, analyst mentions, review platform listings (G2, Capterra, TrustRadius), and industry publication references all contribute to the earned citation layer. A brand that publishes only on its own domain and measures only its own visibility score will miss the third-party corroboration that AI engines use as evidence.

Statistical Rigor: How Much Sampling Is Enough #

AI responses are stochastic — the same prompt produces different answers across sessions. This means single-prompt checks are unreliable. The field is converging on minimum standards:

  • Maximus Labs: 30 runs per query per platform, 95% confidence intervals
  • Yolando: 60–100 executions per prompt for statistical significance
  • Machine Relations Index: evidence floor of 10 observations across 7+ distinct dates before publishing a citation rate

The practical implication: tools offering single-shot audits (HubSpot AEO Grader, free-tier checkers) are useful for initial direction-setting but should not be used to make investment decisions. Continuous measurement with sufficient sampling is required for reliable trend detection. A 5-percentage-point drop in visibility score based on 10 prompt runs may be noise. The same drop sustained across 100 runs over 4 weeks is a signal.

Measurement Cadence: Weekly, Monthly, and Quarterly Reporting #

Yolando's measurement framework provides a practical cadence that aligns with how AI visibility data stabilizes:

Weekly: Run your core prompt library (30–50 prompts) across all tracked engines. Record visibility rate, share of voice, and citation presence at the prompt level. Flag any platform where visibility dropped more than 5 percentage points for investigation.

Monthly: Aggregate weekly data into trend reports. Calculate citation share, prominence scores, and sentiment themes across the competitive set. Compare source-level citation rates if available. Identify which content assets are being cited and which have dropped.

Quarterly: Correlate AI visibility metrics with pipeline data. Review whether content investments produced citation or mention gains. Adjust the prompt library to reflect new buyer language, competitive entrants, and emerging query patterns. Decommission prompts that no longer match real buyer behavior.

Post-publish: After every significant content publish, re-run relevant prompts within 1–2 weeks. This creates a direct feedback loop between content investment and AI visibility outcomes, answering the question editorial teams actually need answered: did this piece move the needle?

Machine Relations and AI Visibility Measurement #

Machine Relations is the discipline of managing how AI systems discover, evaluate, and represent a brand. Measurement is how you prove that relationship exists.

The prompt-monitoring tools listed above measure the buyer-facing output — what AI says about your brand. Source-level citation measurement tracks the structural input — what AI uses as evidence. Both are Machine Relations metrics because both track the relationship between a brand's content and AI system behavior.

The measurement gap most B2B brands face is not a tool gap. It is a framework gap. They monitor brand mentions without understanding what drives those mentions. They track visibility scores without knowing which source assets produce the citations that sustain visibility. They measure the dashboard without measuring the engine.

A complete AI visibility measurement framework layers prompt-based brand monitoring (the demand signal) with source-level citation analysis (the supply signal). When visibility drops, the framework shows whether the brand lost mentions (a positioning problem) or lost citations (a source architecture problem). These are different problems with different solutions.

FAQ #

What is the best free AI visibility measurement tool for a baseline check? #

HubSpot AEO Grader runs 100 test queries across ChatGPT, Perplexity, and Gemini at no cost. TurboAudit offers a free tier with 250+ AI checks. Both provide a directional baseline — sufficient for knowing whether AI engines mention your brand at all, not precise enough for trend tracking or investment decisions.

How is citation rate different from visibility score? #

Visibility score measures whether a brand is mentioned in AI responses to a set of prompts. Citation rate measures how often a specific source domain is cited as evidence across AI answer runs. A brand can score high on visibility (mentioned often) but low on citation rate (content rarely used as a source). The metrics track different events and require different measurement infrastructure.

How many AI engines should a B2B brand track? #

At minimum, track ChatGPT, Google AI Overviews, Perplexity, and Gemini — these represent the highest-volume buyer-facing AI surfaces in mid-2026. Add Claude, Microsoft Copilot, and Google AI Mode for fuller coverage. Profound tracks 11 platforms; Cognizo covers 10 including Grok, Meta AI, and DeepSeek. Single-engine tracking produces blind spots because citation behavior varies across engines.

How often should AI visibility be measured? #

Weekly prompt runs provide reliable trend data when using 30+ executions per prompt per platform. Monthly aggregation supports executive reporting. Quarterly reviews connect visibility data to pipeline outcomes. Single-shot audits are useful for initial baselines but not for ongoing measurement. The Machine Relations Index requires 10 observations across 7+ distinct dates before publishing any citation rate, reflecting the minimum data volume needed for stable measurement.

Do most websites block AI crawlers? #

The Otterly AI Citations Report 2026 found 73% of websites block AI crawlers through robots.txt or CDN rules. Before investing in AI visibility measurement, verify that GPTBot, PerplexityBot, ClaudeBot, and OAI-SearchBot can access your content. A 0% visibility score may indicate an access problem, not a content or brand problem.

Last updated: July 30, 2026. Methodology: Machine Relations Index v2 citation rate data (6 engines, 6,020+ domains, evidence floor of 10 observations across 7+ distinct run dates). Tool pricing and engine coverage verified mid-2026.