Google's AI optimization guide tells publishers how to appear in Google's AI features. It is correct about the foundation: quality content, clear structure, and crawlability remain prerequisites. But because the guide addresses only Google, it structurally cannot measure or explain how citation authority works across the six AI engines that now retrieve and cite web sources — ChatGPT, Perplexity, Gemini, Claude, Google AI Mode, and Google AI Overviews.
Machine Relations Index (MRI) data across 6,020 domains and 17,540 source events shows that cross-engine citation behavior diverges sharply from single-engine optimization logic. For a top-cited domain like G2, Google's own AI features (AI Mode + AI Overviews) account for only 31% of its total citations — the remaining 69% come from Gemini, Perplexity, ChatGPT, and Claude. Sources that rank well in Google's AI features are not always the same sources that ChatGPT, Perplexity, or Claude select. Understanding why requires a framework that measures all six engines simultaneously.
What Google's guide confirms #
Google's guide establishes three principles that MRI data independently validates across all six engines:
Content quality is non-negotiable. Google states publishers should create "unique, compelling, and useful content" with distinctive viewpoints. MRI measurement confirms this: among 6,020 tracked domains, sources with original research, proprietary data, or expert methodology earn higher answer-engine citation rates than commodity aggregators regardless of which engine retrieves them.
Technical crawlability is the floor. The guide confirms that Google's AI features use retrieval-augmented generation (RAG) built on "core Search ranking systems" that require content to be "crawlable" and indexed. Every AI engine with a web crawler — Perplexity, ChatGPT (via GPTBot), Gemini, Claude (via ClaudeBot) — shares this constraint. Blocking any crawler removes a source from that engine's retrieval pool entirely.
Structure helps extraction. Google recommends clear headings, sections, and organized content. MRI position-quality data confirms this mechanically: sources with structured H2/H3 hierarchies and direct answers within the opening 200 words appear in higher citation positions across all engines, not just Google.
Where the guide structurally cannot see #
The guide's limitation is not what it says wrong — it is what a single-engine framework cannot address.
Cross-engine citation divergence #
MRI measurement across the same 30-day window reveals that engines disagree substantially on which sources to cite for identical queries. For the query "AI-powered threat detection for enterprise security," G2 received 145 citations — but the distribution was uneven: Gemini cited G2 in 52 instances, Google AI Mode in 29, Perplexity in 30, Google AI Overviews in 16, ChatGPT in 9, and Claude in 9. Gartner received 130 total citations for the same topic cluster, but with a different engine profile.
A publisher optimizing only for Google's AI features would capture 31% of G2's citations (AI Mode + AI Overviews). The remaining 69% — Gemini at 36%, Perplexity at 21%, ChatGPT at 6%, and Claude at 6% — exist outside Google's optimization framework entirely.
Source role determines citation behavior #
Google's guide does not address source roles. MRI taxonomy classifies every tracked domain by its function in the information supply chain: market database, analyst research, wire distribution, news media, vendor documentation, or community platform.
This classification predicts citation behavior more reliably than any content-level optimization:
| Source Role | Example Domain | 30-day Citations | Engines Citing | Verticals |
|---|---|---|---|---|
| Market database | G2 | 145 | 6/6 | 10 |
| Market database | Crunchbase | 81 | 6/6 | 9 |
| Analyst research | Gartner | 130 | 6/6 | 10 |
| Analyst research | Forbes | 65 | 6/6 | 9 |
| Wire distribution | PR Newswire | 35 | 6/6 | 9 |
| Analyst research | Deloitte | 50 | 6/6 | 8 |
Market databases and analyst research sources dominate citation counts not because their content is better structured for AI extraction, but because their structural role in the information supply chain makes them the default retrieval target for factual queries. Google's guide cannot capture this because it treats all publishers as equivalent optimizers.
Cross-category citation breadth as a predictor #
Google's guide addresses no concept equivalent to cross-category citation breadth — whether a source earns citations across many distinct industry categories. MRI data shows this breadth is one of the strongest predictors of citation authority:
Sources cited across many industry categories (cybersecurity, enterprise AI, fintech, healthtech, HR tech, infrastructure/devtools, and others) sustain higher answer-engine citation rates than sources concentrated in one. G2 earns citations across a wide category spread at high (A) confidence; Deloitte spans fewer.
The mechanism: AI engines retrieve sources that have demonstrated reliability across diverse query contexts. A market database cited in both cybersecurity procurement queries and HR tech comparisons signals broad retrieval value that no single-vertical publisher can replicate through content optimization alone.
The llms.txt disagreement #
Google's guide explicitly dismisses llms.txt files: they receive "no preferential treatment" in Google Search. This is accurate for Google. But llms.txt was designed for direct LLM consumption, not search-engine crawling. Perplexity, Claude, and ChatGPT each maintain independent crawling and retrieval infrastructure where machine-readable metadata can influence source selection differently than Google's ranking systems process it.
MRI does not claim llms.txt causes citations. But a single-engine guide that dismisses a cross-engine signal without measuring its effect on the other five engines leaves publishers with an incomplete picture.
The earned media blind spot #
Google's guide warns against "seeking inauthentic mentions across the web." MRI data draws a sharper distinction: earned media placements — coverage from journalists, analysts, and independent reviewers — correlate with higher citation rates across all six engines. Manufactured mentions do not.
The difference is not optimization technique. It is whether the mention represents real third-party validation that AI retrieval systems can trace. Google's guide conflates genuine earned authority with artificial link-building, losing the distinction that cross-engine measurement reveals.
What Machine Relations measures that single-engine frameworks cannot #
The Machine Relations Index measures each tracked source by its answer-engine citation rate, observed across facets that require multi-engine tracking:
- Engine coverage: How many of the six engines cite the source (all six is the ceiling)
- Query and category breadth: How many distinct queries and industry categories trigger citations
- Position: Where the citation appears in the AI response (higher positions indicate retrieval priority)
- Temporal consistency: How stable citations are across observation dates in the measurement window
None of these facets is measurable from a single engine's optimization guide. Each requires simultaneous tracking across all engines to reflect actual citation authority rather than single-engine ranking signals.
How to use both frameworks #
Google's guide and the Machine Relations framework are not competing. They operate at different layers:
| Layer | Google's Guide | Machine Relations |
|---|---|---|
| Scope | Google AI features only | 6 engines simultaneously |
| Optimization target | Content and technical quality | Source authority and citation patterns |
| Measurement | Google Search Console | MRI citation-rate scoring |
| Source classification | None (treats all publishers equally) | Role taxonomy (market DB, analyst, wire, etc.) |
| Category analysis | None | Cross-category citation breadth |
| Earned media | Warns against "inauthentic mentions" | Distinguishes earned authority from manufactured signals |
Publishers should follow Google's technical and content guidance as the foundation. Then use cross-engine measurement to understand whether their source authority extends beyond Google to the engines where buyers, researchers, and AI assistants also retrieve answers.
FAQ #
Does Google's AI optimization guide apply to ChatGPT and Perplexity? #
No. Google's guide explicitly covers Google's AI features — AI Overviews and AI Mode. ChatGPT, Perplexity, Claude, and Gemini operate independent retrieval systems with different source selection criteria. Semrush's analysis identifies this as the guide's critical limitation: it "applies only to the Google ecosystem."
Is llms.txt worth implementing despite Google dismissing it? #
Google confirmed llms.txt receives no preferential treatment in Google Search. However, llms.txt was designed for direct LLM consumption. Its effect on Perplexity, Claude, and ChatGPT retrieval is a separate measurement question that Google's guide does not and cannot address. MRI tracks cross-engine signals; publishers should evaluate llms.txt based on multi-engine data, not a single engine's dismissal.
What is cross-category citation breadth and why does it predict citation authority? #
Cross-category citation breadth measures how many distinct industry categories (cybersecurity, fintech, healthtech, HR tech, etc.) cite a source. In MRI data, sources cited across many categories sustain higher answer-engine citation rates than sources concentrated in a few. The mechanism: AI engines prioritize sources that demonstrate reliability across diverse contexts, not sources optimized for a single topic cluster.
How does Machine Relations Index differ from traditional SEO metrics? #
Traditional SEO metrics measure ranking position, traffic, and backlinks within a single search engine. MRI measures citation authority across six AI engines simultaneously, scoring sources by their answer-engine citation rate across engines, queries, categories, and observation windows. A source can rank well in Google and be invisible to ChatGPT — MRI captures both.