The GEO Measurement Protocol
How to establish a defensible baseline for your brand’s visibility in AI answer engines — and how to tell a real measurement from a dashboard number.
What is a GEO baseline?#
A GEO baseline is a fixed, timestamped record of how often and how favourably AI answer engines mention a brand, captured across a controlled prompt set before any optimisation work begins. Without one, any later improvement cannot be distinguished from normal answer variability, model updates, or index changes.
A baseline is valid only if four conditions hold: the prompt wording is fixed and recorded; each prompt is run multiple times; the engine, product mode, location, and timestamp are recorded for every run; and the metric denominator is stated explicitly.
Why measurement is the hard part#
Most brands attempting GEO are optimising against a number they cannot reproduce.
The reason is that AI answers are not stable objects. A 2026 arXiv study found identical prompts produced distinct outputs roughly 25% of the time on GPT-4o-mini and roughly 10% on Llama 3.1 8B under its test conditions. A separate 2026 analysis found that query wording alone accounted for approximately 26.5% of response variance, while brand identity accounted for approximately 1.5%.
Google states directly that AI Overviews and AI Mode may use different models and techniques, so responses and links can vary between the two surfaces. Google also documents query fan-out — the system runs multiple related searches across subtopics rather than one query.
The practical consequence: a single run of a single prompt tells you almost nothing. Two runs a week apart showing different results is the expected behaviour of the system, not evidence that anything changed.
This is why the first deliverable of any serious GEO engagement is not a fix. It is a baseline.
What is actually at stake#
The measurement problem matters because the traffic shift is already measurable.
Pew Research Center found users clicked a traditional Google result in 8% of visits where an AI Overview appeared, compared with 15% of visits where none appeared, across 68,879 searches from 900 US adults. Ahrefs reported that AI Overview presence was associated with a 58% lower average click-through rate for the top-ranking page. Seer Interactive reported organic CTR falling from 1.76% to 0.61% across its tracked AI Overview queries.
Scale figures, all company-reported rather than independently audited: OpenAI said ChatGPT was on track for 700 million weekly active users in August 2025. Perplexity’s CEO stated 780 million queries in May 2025. Google reported AI Overviews reaching more than 1.5 billion monthly users across more than 200 countries in May 2025.
Prevalence, however, is volatile and dataset-dependent. Semrush tracked AI Overviews on 6.49% of its keywords in January 2025, 24.61% in July 2025, and 15.69% in November 2025. Any figure presented as “the percentage of Google searches with AI Overviews” without naming the panel, country, device, and date should be treated as unreliable.
What this means for a brand: the click is worth less and the citation is worth more. That reallocation is measurable. Whether it converts is not yet proven — see section 07.
Define the unit before you measure it#
The most common measurement failure is not sampling. It is measuring four different things and reporting them as one number.
The IAB’s Measuring Visibility in the AI Era framework, published 3 August 2026, separates AI visibility into four layers. We use its vocabulary because a shared vocabulary is more useful than a proprietary one.
Layered on top, we separate two things most tools conflate:
- Mention — the brand name appears in the generated text.
- Citation — the brand’s domain appears in the linked source set.
These diverge constantly. A brand can be recommended without its site being cited, and cited as a source without being recommended. The 2026 academic measurement framework makes a related distinction between citation selection — being chosen as a source — and citation absorption — the answer actually using what that source said.
A visibility score that collapses all of this into one percentage is directional. It is not decision-grade. The IAB framework itself draws that line.
The protocol#
PHASE 01Scope Lock
Build the prompt set before touching the site, then freeze it. Two categories, kept separate throughout.
Identity prompts test whether the engine knows the brand exists and describes it accurately. What is the brand, where is it located, what does it offer, who runs it.
Commercial prompts test whether the engine recommends the brand when the buyer never names it. Best provider in a category, alternatives to a competitor, how to choose within the category.
Identity prompts measure whether you exist to the model. Commercial prompts measure whether you are chosen. They are different problems with different fixes, and averaging them produces a number that describes neither.
Every prompt is scored against five binary criteria, fixed before measurement begins:
C5 is the criterion no public framework includes, and in practice it fails more often than the other four combined. An engine that confidently describes the wrong company is not a visibility problem. It is a factual one.
Published guidance clusters around 20–50 prompts, or at least 50 with three runs each. These are recommendations, not validated power calculations. The required sample depends on your baseline citation rate — a set sized for a 50% mention rate is badly underpowered for a 2% rate.
Our standing baseline is 296 captures — 144 identity and 152 commercial — which produces stable per-engine and per-category rates at the citation frequencies we observe in mid-market brands.
PHASE 02Environment Control
Every variable that is not the prompt must be fixed and recorded:
- Fresh, logged-out sessions per run — no chat history, no personalisation, no memory
- Fixed geographic location, stated explicitly
- Fixed language, with multilingual brands measured per language rather than pooled
- Product mode named exactly — ChatGPT Search is not the ChatGPT API, and AI Mode is not AI Overviews
- Timestamp on every capture
Yext explicitly documents that citations returned by APIs may differ from those shown in the consumer application. This is the most under-disclosed limitation in commercial AI visibility tooling. We capture from consumer interfaces, not APIs, because the consumer interface is what a buyer actually sees.
PHASE 03Capture
Every run produces one row and one screenshot. The screenshot is not decoration — it is the evidence that the row is real, and the record that survives when the engine produces something different next week. A standard baseline produces 440 screenshots across the measurement window.
Responses are recorded in full. Nothing is summarised, trimmed, or paraphrased at capture time, because scoring rubrics get revised and a summarised response cannot be re-scored.
| FIELD | WHY IT IS RECORDED |
|---|---|
| Exact prompt text | Wording drives ~26.5% of response variance |
| Engine and product mode | Retrieval systems differ; AI Mode ≠ AI Overviews |
| Model version, where exposed | Model updates shift outputs independently of your site |
| Timestamp | Enables volatility analysis across the window |
| Location and language | Local intent changes source selection |
| Session state | Logged out, fresh, no history |
| Full response text | Required for re-scoring under a revised rubric |
| Brand mentioned | Presence |
| Brand domain cited | Citation — tracked separately from mention |
| Position of mention | Prominence |
| Mention type | Recommended, listed, compared, warned against |
| Sentiment | Portrayal |
| Factual accuracy | Whether claims about the brand are correct |
| All cited URLs | Reveals which third-party sources the engine trusts |
| Competitors named | Share of voice |
| Screenshot reference | Evidence and audit trail |
PHASE 04Dominance Scoring
Each prompt runs three times per engine. The majority result becomes the dominant score for that prompt and engine. Where the three runs disagree, the disagreement is recorded as volatility rather than averaged away.
This is what converts a snapshot into a rate with an error range. A brand appearing in 3 of 12 runs has a 25% mention rate with real uncertainty attached — a meaningfully different statement from “we appeared in ChatGPT.”
Non-answers score zero on all applicable criteria. A refusal, an empty result, or a redirect to a search page is not a neutral outcome; from the buyer’s position it is identical to absence.
PHASE 05Access Verification
Retrieval access is a precondition, not an optimisation. Verify each control separately, because the crawler documentation is genuinely counter-intuitive.
| ENGINE | SEARCH INCLUSION | TRAINING / OTHER |
|---|---|---|
| ChatGPT | OAI-SearchBot | GPTBot (model improvement), ChatGPT-User (user-triggered fetch) |
| Perplexity | PerplexityBot | Perplexity-User (user-requested fetch, not general crawl) |
Googlebot plus index and snippet eligibility | Google-Extended (certain AI training and grounding) | |
| Claude | Claude-SearchBot | ClaudeBot (model training) |
| Copilot | Bing indexing | — |
The most common technical error we find: allowing GPTBot and assuming the site is eligible for ChatGPT Search. OpenAI documents these as separate crawlers with separate purposes. Blocking OAI-SearchBot removes the site from ChatGPT search answers regardless of GPTBot status.
Google states that to be eligible as a supporting link, a page must be indexed and eligible to appear with a snippet — and that there are no additional technical requirements specific to AI Overviews or AI Mode.
PHASE 06Baseline Report
Metrics are reported with their denominators stated. Without the denominator, cross-study and cross-tool comparison is meaningless.
| METRIC | DEFINITION |
|---|---|
| Mention rate | Runs where the brand name appears ÷ valid runs |
| Citation rate | Runs where the brand domain appears in sources ÷ valid runs |
| Share of voice | Brand mentions ÷ all brand mentions in the same run set |
| Accuracy rate | Runs with factually correct brand claims ÷ runs mentioning the brand |
| Volatility | Repeat runs producing a different outcome for the same prompt |
Volatility is reported alongside every rate. A 40% mention rate with 5% volatility and a 40% mention rate with 30% volatility are not the same finding, and only one of them supports a decision.
The report also carries an error register — every factual mistake an engine made about the brand, listed individually with the engine, language, and date it occurred. That register, not the score, is usually what a client acts on first.
Reading the results#
A completed baseline typically resolves into one of four diagnoses.
On the third case: a Yext analysis of 6.8 million citations across 1.6 million responses on ChatGPT, Gemini, and Perplexity found approximately 86% came from brand-controlled or brand-associated sources — though this is a vendor study drawn from its own client base and has not been independently replicated.
Misrepresentation is the most urgent diagnosis and the one dashboards miss most often, because a mention is scored as a positive regardless of what it says. Increasing the visibility of a wrong answer makes the problem worse, not better.
Each diagnosis leads to a different intervention. This is the entire argument for baselining before optimising.
On the relationship to search rankings#
Ranking first in Google does not guarantee inclusion in AI answers, and the overlap studies are frequently misquoted because they use different denominators.
seoClarity reported approximately 32% of AI Overview citations overlapped with Google’s top ten, and separately that approximately 90% of AI Overviews contained at least one top-ten result. Ahrefs reported approximately 37.9% of cited URLs also ranked in the organic top ten, across 863,000 keywords and roughly 4 million AI Overview URLs.
These are not contradictory. The first pair measures citations; the second measures whole Overviews. Comparing them requires matching the denominator, country, query set, date, and definition of overlap.
Google’s own documentation notes AI Overviews can surface a wider and more diverse set of links than classic Search — which implies ranking position is not the only source-selection input.
Reading: organic ranking is correlated with citation and is worth having. It is neither necessary nor sufficient, and the overlap studies do not establish causation in either direction.