# What the IAB's Decision-Grade Standard Means for AI Visibility Measurement

The IAB defined decision-grade AI visibility measurement. Here is what the standard requires, what most tools miss, and what 10,487 observed answer-engine runs reveal about meeting the bar.

Canonical URL: https://machinerelations.ai/research/iab-decision-grade-standard-mri-citation-rates-2026
Published: 2026-08-03
Research type: Index Analysis
Tags: ai-visibility, measurement, mri

## Source Body

The IAB released [Measuring Visibility in the AI Era](https://iab.com/guidelines/measuring-visibility-in-the-ai-era) on August 3, 2026, giving brands, agencies, and publishers a shared framework for evaluating how they appear inside AI-powered discovery. The most useful contribution is not the framework itself but a distinction buried inside it: directional measurement versus decision-grade measurement. That distinction determines whether an AI visibility number is worth acting on.

## The IAB separates directional from decision-grade measurement

The IAB framework organizes AI visibility into [four categories](https://www.prnewswire.com/news-releases/iab-releases-measuring-visibility-in-the-ai-era-to-help-brands-publishers-and-agencies-navigate-ai-powered-discovery-302840611.html) — presence, prominence, portrayal, and persuasion — and then introduces a quality bar that most current tools do not meet.

**Directional measurement** identifies patterns and trends. It tells a brand that something is happening. It does not tell the brand enough to make a budget decision, change a content strategy, or evaluate a vendor.

**Decision-grade measurement** requires six properties:

| Requirement | What it means operationally |
|---|---|
| Query volume | Enough distinct prompts to represent the buyer question space |
| Sample size | Enough observed runs per segment to produce stable rates |
| Prompt coverage | Queries spanning the brand's actual category and competitor set |
| Testing cadence | Regular observation over time, not a single snapshot |
| Reproducibility | Methodology that produces consistent results across runs |
| Platform coverage | Multiple AI engines measured, not a single provider |

Over [20 companies now sell AI visibility measurement tools](https://www.marketingdive.com/news/iab-shares-playbook-for-measuring-brand-visibility-in-ai-powered-platforms/826830/), each using a different methodology. Caroline Gigerich, VP of AI at the IAB, framed the problem as an industry-wide inconsistency that prevents brands from distinguishing reliable signals from noise. The framework does not name which tools meet the decision-grade bar. It gives the industry criteria to evaluate them.

## What 10,487 observed runs reveal about meeting the bar

The [Machine Relations Index](https://machinerelations.ai/glossary/machine-relations-index) v2 measures source-segment citation rates across six answer engines — ChatGPT, Claude, Gemini, Google AI Mode, Google AI Overviews, and Perplexity — on a daily cadence. Over an 81-day window ending August 3, 2026, the index observed 10,487 answer-engine runs covering 17,145 distinct domains and 88,747 source events.

Mapping MRI v2 against the IAB's decision-grade requirements:

| IAB requirement | MRI v2 implementation |
|---|---|
| Query volume | Prompts span 10 B2B verticals with multiple buyer question types per vertical |
| Sample size | Evidence floor requires 10+ observations across 7+ distinct run dates before publishing a citation rate |
| Prompt coverage | Each segment pairs a subject category with a buyer question shape |
| Testing cadence | Daily observation over 81 consecutive days |
| Reproducibility | Citation rates are computed as runs-cited divided by runs-observed per segment; methodology is published |
| Platform coverage | Six engines measured simultaneously |

Not all domains clear the evidence floor. Domains below 10 observations across 7 run dates are classified as "collecting" rather than scored. This prevents thin data from producing misleading rates — exactly the kind of statistical discipline the IAB framework calls for but does not enforce.

## Citation rates are not visibility scores

The IAB framework's first category, presence, asks whether a brand appears in AI responses at all. Most current tools answer this with a binary yes/no or a composite score. The MRI takes a different approach: it reports the rate at which a domain is cited across observed runs.

Forbes, for example, carries a 4.18% citation rate — cited in 438 of 10,487 runs, across all six engines, on 66 of 81 observed days. That number is more useful than a "visibility score" because it has a denominator. A brand team can compare rates across domains, across time, and across segments without wondering what formula produced the number.

The distinction matters because rates degrade gracefully. A [comparison of AI visibility tools in 2026](https://www.stork.ai/blog/best-ai-visibility-tools-2026) found that vendors use different scoring formulas, making cross-platform comparison difficult. A score of 79 means nothing without knowing the scale, the weighting, and the threshold. A citation rate of 4.18% means the domain appeared in roughly 1 of every 24 observed answer runs. If the rate drops next month, the team knows the magnitude of the change without recalibrating a proprietary index.

This is also why the [difference between LLM visibility and LLM acquisition](https://kissmetrics.io/blog/llm-visibility-vs-llm-acquisition) matters operationally. Being mentioned by an AI engine is a presence signal. Converting that mention into a site visit, pipeline entry, or purchase decision is an acquisition outcome. The IAB framework captures this split in its portrayal-to-persuasion progression, but most current measurement sits entirely in the presence layer.

## Portrayal and persuasion remain unmeasured at scale

The IAB's third and fourth categories — portrayal and persuasion — describe measurement problems that no public dataset has solved operationally.

Portrayal asks whether the brand is described accurately. That requires comparing what the AI said against what the brand considers true — a factual-alignment check that demands both brand-specific ground truth and natural-language comparison at scale. Current approaches either require manual review (accurate but unscalable) or use LLMs to evaluate LLM output (scalable but circular). Research on [rank stability in AI visibility measurement](https://arxiv.org/html/2607.10341v1) treats this as a structural sufficiency question — even presence data is unstable without enough observations, and portrayal data is an order of magnitude harder to stabilize.

Persuasion asks whether the AI response moved the user toward action. Measuring post-citation behavior requires tracking what happens after a user reads an AI-generated answer — clicks, conversions, or preference changes — and attributing them to a specific mention in a specific response. Research on [uncertainty in generative search measurement](https://arxiv.org/pdf/2603.08924) frames the entire AI answer measurement problem as statistical rather than deterministic, which makes persuasion attribution — already hard in traditional media — even less tractable. No measurement system currently connects these events reliably across engines.

The operational implication is that presence and prominence measurement is solvable now. Portrayal and persuasion measurement is a research problem. Brands that wait for a single tool covering all four categories will wait longer than brands that start measuring what can be measured today.

## What the framework does not address

The IAB framework is a vocabulary contribution, not an enforcement mechanism. It does not:

- **Certify tools.** No vendor receives a decision-grade stamp from the IAB. The criteria exist for buyers to evaluate vendors themselves.
- **Specify sampling methodology.** Two tools can both claim query volume and platform coverage while measuring fundamentally different prompt sets with different observation frequencies.
- **Address source architecture.** The framework measures outcomes (does the brand appear?) but does not connect measurement to the source-layer work that changes outcomes. [Tracking AI citation patterns across engines](https://www.searchenginejournal.com/comparison-of-ai-citation-patterns-offers-strategic-seo-insights/573327/) reveals that different engines cite different source types for the same query category, which means a brand can track its presence rate daily and still have no clear path to improving it without understanding which sources each engine actually trusts.

This last gap is where [Machine Relations](https://machinerelations.ai/glossary/machine-relations) as a discipline goes beyond measurement. The MRI does not only report whether a brand appears. It reports which source domains AI engines cite, at what rates, across which categories and question types. That source-level data connects a measurement result to an actionable publishing or earned-media decision. The [IETF AI Visibility Lifecycle Framework](https://datatracker.ietf.org/doc/draft-lynch-ai-visibility-lifecycle) draft, which has progressed independently of the IAB, approaches the same problem from an infrastructure perspective — defining lifecycle stages for how content moves from creation through AI retrieval to citation.

## How to evaluate AI visibility tools against the IAB bar

For brands choosing a measurement provider, the IAB framework provides a practical checklist. Ask the vendor:

1. **How many distinct prompts do you observe?** A tool measuring 50 queries provides directional signal. A tool measuring thousands across structured buyer categories provides decision-grade coverage.
2. **How many engines do you measure?** Single-engine measurement captures one platform's behavior. Cross-engine measurement reveals which sources carry authority broadly versus on a single provider.
3. **What is your observation cadence?** Monthly snapshots miss the volatility documented in [repeated AI search visibility measurement research](https://arxiv.org/abs/2604.07585). Daily or weekly cadence produces stable distributions. [Tinuiti's tracking of AI citation patterns](https://tinuiti.com/blog/search/ai-citations/) shows that citation behavior shifts across engine updates and seasonal cycles, reinforcing the need for continuous rather than periodic observation.
4. **Do you publish your methodology?** If the scoring formula is proprietary and undisclosed, the buyer cannot distinguish signal from noise.
5. **Do you report rates or scores?** Rates have denominators and degrade gracefully. Composite scores require the buyer to trust a weighting scheme they cannot inspect.

## FAQ

### What is IAB's Measuring Visibility in the AI Era?

[Measuring Visibility in the AI Era](https://iab.com/guidelines/measuring-visibility-in-the-ai-era) is an IAB framework released August 3, 2026. It defines four measurement categories (presence, prominence, portrayal, persuasion) and separates directional measurement from decision-grade measurement based on query volume, sample size, cadence, reproducibility, and platform coverage.

### What is the difference between directional and decision-grade AI visibility measurement?

Directional measurement identifies trends and patterns but is not rigorous enough for budget allocation or vendor evaluation. Decision-grade measurement meets standards for query volume, sample size, prompt coverage, testing cadence, reproducibility, and platform coverage — producing data that supports operational decisions.

### How does the Machine Relations Index measure AI visibility?

The [Machine Relations Index](https://machinerelations.ai/glossary/machine-relations-index) v2 reports citation rates — the fraction of observed answer-engine runs in which a domain is cited — across six engines on a daily cadence. Domains must clear an evidence floor of 10+ observations across 7+ distinct run dates before a rate is published, preventing thin data from creating false signals.

### Which AI visibility measurement categories can be measured today?

Presence and prominence are operationally measurable now through citation-rate tracking across engines. Portrayal (accuracy of brand descriptions) and persuasion (post-citation user behavior) require ground-truth comparison and cross-platform attribution that no public measurement system has solved at scale.

## Attribution

This research is published by Machine Relations Research, the research program of machinerelations.ai — the public research and standards initiative that publishes the glossary, research, evidence, and measurements for the Machine Relations discipline. Provenance and editorial standards: https://machinerelations.ai/about

## Machine-readable related links

### Related concepts

- [Machine Relations Index (MRI)](https://machinerelations.ai/glossary/machine-relations-index)
- [Machine Relations (MR)](https://machinerelations.ai/glossary/machine-relations)
- [AI Visibility](https://machinerelations.ai/glossary/ai-visibility)
- [AI Share of Voice (AI SOV)](https://machinerelations.ai/glossary/ai-share-of-voice)

### Supporting research

- [How to Get Cited by ChatGPT: What Source Selection Data Actually Shows](https://machinerelations.ai/research/how-to-get-cited-by-chatgpt-mri-source-selection-data-2026)
- [What Is AI Share of Voice? Definition, Formula, and Measurement Framework (2026)](https://machinerelations.ai/research/what-is-ai-share-of-voice)
- [Google Search Console AI Performance Reports: What They Measure, What They Miss, and Why It Matters for Citation Architecture](https://machinerelations.ai/research/google-search-console-ai-performance-reports-citation-architecture)
- [Semrush AI Visibility Index: 126M Prompts, Citation Gap](https://machinerelations.ai/research/semrush-ai-visibility-index-126-million-prompts-citation-authority-2026)

### Framework context

- [Machine Relations Stack](https://machinerelations.ai/stack)
- [Evidence Base](https://machinerelations.ai/evidence)
