DeviLabs Book a call
GEO Services Pricing Work Case studies About Blog Book a call
info@devilab.eu · Wassenaar, NL
Contents
01Why measurement is the hard part 02What is actually at stake 03Define the unit before you measure it 04The protocol 05Reading the results 06Relationship to search rankings 07What this protocol does not claim 08Field notes Frequently asked References
DeviLabs · Technical Protocol

The GEO Measurement Protocol

How to establish a defensible baseline for your brand’s visibility in AI answer engines — and how to tell a real measurement from a dashboard number.

Dominykas Jankauskas, Founder of DeviLabs
By Dominykas Jankauskas · Founder, DeviLabs
Published 8 August 2026 · Version 1.0 · Next review November 2026

What is a GEO baseline?#

A GEO baseline is a fixed, timestamped record of how often and how favourably AI answer engines mention a brand, captured across a controlled prompt set before any optimisation work begins. Without one, any later improvement cannot be distinguished from normal answer variability, model updates, or index changes.

A baseline is valid only if four conditions hold: the prompt wording is fixed and recorded; each prompt is run multiple times; the engine, product mode, location, and timestamp are recorded for every run; and the metric denominator is stated explicitly.

01

Why measurement is the hard part#

Most brands attempting GEO are optimising against a number they cannot reproduce.

The reason is that AI answers are not stable objects. A 2026 arXiv study found identical prompts produced distinct outputs roughly 25% of the time on GPT-4o-mini and roughly 10% on Llama 3.1 8B under its test conditions. A separate 2026 analysis found that query wording alone accounted for approximately 26.5% of response variance, while brand identity accounted for approximately 1.5%.

Google states directly that AI Overviews and AI Mode may use different models and techniques, so responses and links can vary between the two surfaces. Google also documents query fan-out — the system runs multiple related searches across subtopics rather than one query.

The practical consequence: a single run of a single prompt tells you almost nothing. Two runs a week apart showing different results is the expected behaviour of the system, not evidence that anything changed.

This is why the first deliverable of any serious GEO engagement is not a fix. It is a baseline.

Diagram 1 · Volatility
The same fixed prompt produces three different runsA single prompt node on the left connects to three separate result sets on the right. Run one lists two competitors and no mention of the brand. Run two mentions the brand in third position with no citation. Run three cites the brand domain as a source. The same prompt produces three different outputs. PROMPT · FIXED best dog boarding wassenaar RUN 01 Two competitors named mention: no · citation: no RUN 02 Brand third in a list of five mention: yes · citation: no RUN 03 Brand domain in source set mention: yes · citation: yes
The same prompt, three runs. This is normal engine behaviour, not a change in your visibility.
02

What is actually at stake#

The measurement problem matters because the traffic shift is already measurable.

Pew Research Center found users clicked a traditional Google result in 8% of visits where an AI Overview appeared, compared with 15% of visits where none appeared, across 68,879 searches from 900 US adults. Ahrefs reported that AI Overview presence was associated with a 58% lower average click-through rate for the top-ranking page. Seer Interactive reported organic CTR falling from 1.76% to 0.61% across its tracked AI Overview queries.

8% vs 15%
Click-through on traditional results, with vs without an AI Overview
Pew Research Center, 68,879 searches, July 2025
58%
Lower average CTR for the top-ranking page when an AI Overview is present
Ahrefs, February 2026
6.49% → 24.61% → 15.69%
AI Overview prevalence across tracked keywords, Jan / Jul / Nov 2025
Semrush

Scale figures, all company-reported rather than independently audited: OpenAI said ChatGPT was on track for 700 million weekly active users in August 2025. Perplexity’s CEO stated 780 million queries in May 2025. Google reported AI Overviews reaching more than 1.5 billion monthly users across more than 200 countries in May 2025.

Prevalence, however, is volatile and dataset-dependent. Semrush tracked AI Overviews on 6.49% of its keywords in January 2025, 24.61% in July 2025, and 15.69% in November 2025. Any figure presented as “the percentage of Google searches with AI Overviews” without naming the panel, country, device, and date should be treated as unreliable.

What this means for a brand: the click is worth less and the citation is worth more. That reallocation is measurable. Whether it converts is not yet proven — see section 07.

03

Define the unit before you measure it#

The most common measurement failure is not sampling. It is measuring four different things and reporting them as one number.

The IAB’s Measuring Visibility in the AI Era framework, published 3 August 2026, separates AI visibility into four layers. We use its vocabulary because a shared vocabulary is more useful than a proprietary one.

Diagram 2 · The four-layer stack
The four-layer AI-visibility stack: Presence, Prominence, Portrayal, Persuasion — each with its own metric.

Layered on top, we separate two things most tools conflate:

  • Mention — the brand name appears in the generated text.
  • Citation — the brand’s domain appears in the linked source set.

These diverge constantly. A brand can be recommended without its site being cited, and cited as a source without being recommended. The 2026 academic measurement framework makes a related distinction between citation selection — being chosen as a source — and citation absorption — the answer actually using what that source said.

A visibility score that collapses all of this into one percentage is directional. It is not decision-grade. The IAB framework itself draws that line.

Diagram 3 · Mention vs citation
Mention versus citation are distinct setsTwo overlapping circles. The left circle is labelled MENTIONED, named in the answer text. The right circle is labelled CITED, domain appears in sources. The overlap is labelled BOTH. MENTIONED named in the answer text CITED domain appears in sources BOTH
Different problems. Different fixes. Most tools report one number for both.
04

The protocol#

Diagram 4 · Six-phase pipeline
30-day measurement window
The six-phase baseline pipeline, run across a 30-day measurement window.

PHASE 01Scope Lock

Build the prompt set before touching the site, then freeze it. Two categories, kept separate throughout.

Identity prompts test whether the engine knows the brand exists and describes it accurately. What is the brand, where is it located, what does it offer, who runs it.

Commercial prompts test whether the engine recommends the brand when the buyer never names it. Best provider in a category, alternatives to a competitor, how to choose within the category.

Identity prompts measure whether you exist to the model. Commercial prompts measure whether you are chosen. They are different problems with different fixes, and averaging them produces a number that describes neither.

Every prompt is scored against five binary criteria, fixed before measurement begins:

C1Correct brand name
C2Correct location
C3Correct core services
C4Correct people
C5Not confused with another entity, group company, or country

C5 is the criterion no public framework includes, and in practice it fails more often than the other four combined. An engine that confidently describes the wrong company is not a visibility problem. It is a factual one.

Published guidance clusters around 20–50 prompts, or at least 50 with three runs each. These are recommendations, not validated power calculations. The required sample depends on your baseline citation rate — a set sized for a 50% mention rate is badly underpowered for a 2% rate.

Our standing baseline is 296 captures — 144 identity and 152 commercial — which produces stable per-engine and per-category rates at the citation frequencies we observe in mid-market brands.

PHASE 02Environment Control

Every variable that is not the prompt must be fixed and recorded:

  • Fresh, logged-out sessions per run — no chat history, no personalisation, no memory
  • Fixed geographic location, stated explicitly
  • Fixed language, with multilingual brands measured per language rather than pooled
  • Product mode named exactly — ChatGPT Search is not the ChatGPT API, and AI Mode is not AI Overviews
  • Timestamp on every capture

Yext explicitly documents that citations returned by APIs may differ from those shown in the consumer application. This is the most under-disclosed limitation in commercial AI visibility tooling. We capture from consumer interfaces, not APIs, because the consumer interface is what a buyer actually sees.

PHASE 03Capture

Every run produces one row and one screenshot. The screenshot is not decoration — it is the evidence that the row is real, and the record that survives when the engine produces something different next week. A standard baseline produces 440 screenshots across the measurement window.

Responses are recorded in full. Nothing is summarised, trimmed, or paraphrased at capture time, because scoring rubrics get revised and a summarised response cannot be re-scored.

Table 1 · Capture fieldsscroll →
Table 1 — Capture fields
FIELD WHY IT IS RECORDED
Exact prompt textWording drives ~26.5% of response variance
Engine and product modeRetrieval systems differ; AI Mode ≠ AI Overviews
Model version, where exposedModel updates shift outputs independently of your site
TimestampEnables volatility analysis across the window
Location and languageLocal intent changes source selection
Session stateLogged out, fresh, no history
Full response textRequired for re-scoring under a revised rubric
Brand mentionedPresence
Brand domain citedCitation — tracked separately from mention
Position of mentionProminence
Mention typeRecommended, listed, compared, warned against
SentimentPortrayal
Factual accuracyWhether claims about the brand are correct
All cited URLsReveals which third-party sources the engine trusts
Competitors namedShare of voice
Screenshot referenceEvidence and audit trail

PHASE 04Dominance Scoring

Each prompt runs three times per engine. The majority result becomes the dominant score for that prompt and engine. Where the three runs disagree, the disagreement is recorded as volatility rather than averaged away.

This is what converts a snapshot into a rate with an error range. A brand appearing in 3 of 12 runs has a 25% mention rate with real uncertainty attached — a meaningfully different statement from “we appeared in ChatGPT.”

Non-answers score zero on all applicable criteria. A refusal, an empty result, or a redirect to a search page is not a neutral outcome; from the buyer’s position it is identical to absence.

PHASE 05Access Verification

Retrieval access is a precondition, not an optimisation. Verify each control separately, because the crawler documentation is genuinely counter-intuitive.

Table 2 · Crawler referencescroll →
Table 2 — Crawler reference
ENGINE SEARCH INCLUSION TRAINING / OTHER
ChatGPTOAI-SearchBotGPTBot (model improvement), ChatGPT-User (user-triggered fetch)
PerplexityPerplexityBotPerplexity-User (user-requested fetch, not general crawl)
GoogleGooglebot plus index and snippet eligibilityGoogle-Extended (certain AI training and grounding)
ClaudeClaude-SearchBotClaudeBot (model training)
CopilotBing indexing

The most common technical error we find: allowing GPTBot and assuming the site is eligible for ChatGPT Search. OpenAI documents these as separate crawlers with separate purposes. Blocking OAI-SearchBot removes the site from ChatGPT search answers regardless of GPTBot status.

Google states that to be eligible as a supporting link, a page must be indexed and eligible to appear with a snippet — and that there are no additional technical requirements specific to AI Overviews or AI Mode.

PHASE 06Baseline Report

Metrics are reported with their denominators stated. Without the denominator, cross-study and cross-tool comparison is meaningless.

Table 3 · Metric definitionsscroll →
Table 3 — Metric definitions
METRIC DEFINITION
Mention rateRuns where the brand name appears ÷ valid runs
Citation rateRuns where the brand domain appears in sources ÷ valid runs
Share of voiceBrand mentions ÷ all brand mentions in the same run set
Accuracy rateRuns with factually correct brand claims ÷ runs mentioning the brand
VolatilityRepeat runs producing a different outcome for the same prompt

Volatility is reported alongside every rate. A 40% mention rate with 5% volatility and a 40% mention rate with 30% volatility are not the same finding, and only one of them supports a decision.

The report also carries an error register — every factual mistake an engine made about the brand, listed individually with the engine, language, and date it occurred. That register, not the score, is usually what a client acts on first.

05

Reading the results#

A completed baseline typically resolves into one of four diagnoses.

Diagram 5 · Four diagnoses
The four diagnoses a completed baseline resolves into — each pointing to a different fix.

On the third case: a Yext analysis of 6.8 million citations across 1.6 million responses on ChatGPT, Gemini, and Perplexity found approximately 86% came from brand-controlled or brand-associated sources — though this is a vendor study drawn from its own client base and has not been independently replicated.

Misrepresentation is the most urgent diagnosis and the one dashboards miss most often, because a mention is scored as a positive regardless of what it says. Increasing the visibility of a wrong answer makes the problem worse, not better.

Each diagnosis leads to a different intervention. This is the entire argument for baselining before optimising.

06

On the relationship to search rankings#

Ranking first in Google does not guarantee inclusion in AI answers, and the overlap studies are frequently misquoted because they use different denominators.

seoClarity reported approximately 32% of AI Overview citations overlapped with Google’s top ten, and separately that approximately 90% of AI Overviews contained at least one top-ten result. Ahrefs reported approximately 37.9% of cited URLs also ranked in the organic top ten, across 863,000 keywords and roughly 4 million AI Overview URLs.

~37.9%
Of AI Overview cited URLs also rank in Google’s organic top ten
Ahrefs, 863,000 keywords, March 2026

These are not contradictory. The first pair measures citations; the second measures whole Overviews. Comparing them requires matching the denominator, country, query set, date, and definition of overlap.

Google’s own documentation notes AI Overviews can surface a wider and more diverse set of links than classic Search — which implies ranking position is not the only source-selection input.

Reading: organic ranking is correlated with citation and is worth having. It is neither necessary nor sufficient, and the overlap studies do not establish causation in either direction.

07

What this protocol does not claim#

We publish this because most GEO material asserts things the evidence does not support. Separating the two is the point of a protocol. The four claims below are the ones this protocol touches directly; we checked twenty-four of them against primary sources and controlled studies in The GEO Claims Audit.

Schema markup causes AI citations.

Google recommends structured data match visible text and states that pages need ordinary Search eligibility. It does not state that schema independently raises citation probability. Vendor experiments reporting large schema effects typically change page structure, content, and internal linking at the same time.

Backlinks, Domain Rating, and E-E-A-T are AI ranking factors.

The foundational GEO paper did not isolate any of these as causal variables. They remain plausible hypotheses, not replicated findings.

AI-referred traffic converts better than organic.

No independently replicated cross-engine study establishes this. ChatGPT referrals can be identified via a URL parameter under OpenAI’s publisher documentation, but referral attribution does not demonstrate the answer caused the conversion.

One tactic works across all engines.

ChatGPT, Perplexity, Google, Claude, and Copilot have documented differences in crawlers, retrieval, and product behaviour. No universal citation-ranking formula is public for any of them.

Citation gains are durable.

No independent longitudinal study establishes the causal half-life of a GEO intervention across engines.

What the foundational research does support: Aggarwal et al., in the paper that introduced the term GEO on 16 November 2023, tested content interventions against a benchmark of roughly 10,000 queries and reported visibility improvements up to approximately 40% for the strongest interventions — adding citations, quotations, statistics, and improving fluency. That figure is a relative, experiment-specific result under controlled conditions. It is not a forecast for a commercial website, and the paper reports that effectiveness varies by domain.

What this protocol therefore is: a measurement instrument. It tells you where you stand and whether something changed. It does not tell you that any given lever caused the change, and we will not claim otherwise in a client report.

08

Field notes#

We run this protocol on our own brands before we run it for anyone else.

Case · LuckyPaws

LuckyPaws — dog walking, daycare, and boarding in Wassenaar, Netherlands. Twelve identity and commercial prompts, four engines, Dutch and English, single run per prompt. A short-form pass, not a full baseline.

Pricing was wrong, not missing. ChatGPT returned specific prices for LuckyPaws services that do not match the published rates on the site. The rates are public, on a page a crawler can reach. The engine produced numbers anyway. A missing price is a gap; a confident wrong price is a figure a customer arrives with.
Entity resolution failed on a name collision. Gemini did not resolve LuckyPaws to the Wassenaar business. This is criterion C5 — not confused with another entity — and it is the criterion that fails most often in our client work. It is also invisible to any tool that scores a mention as a positive without checking which company was described.
The service area was placed incorrectly. Gemini described LuckyPaws as operating somewhere other than where it does. For a local service business, a wrong service area is not a partial answer. It removes the business from consideration for every customer who is actually in range.

None of these are visibility problems. The brand appeared in every case. A mention-rate dashboard would have counted all three as successes.

This is the Misrepresented diagnosis from section 05, found on our own brand, on a site with published pricing and a verified Google Business Profile. Presence was never the issue. Accuracy was.

It is the reason we score factual accuracy separately from presence, and the reason correction precedes amplification in every engagement we run.

Frequently asked#

How long does a baseline take?

Thirty days of capture. The window is not padding — it is what allows volatility to be separated from genuine change. A one-day snapshot cannot distinguish the two.

How many prompts do I actually need?

It depends on your baseline citation rate, not on a universal number. Published guidance ranges from 20–50 prompts upward, but these are recommendations rather than power calculations. A brand appearing in 2% of answers needs a substantially larger sample than one appearing in 50%.

Can I just use an AI visibility tool?

Tools are useful for trend monitoring. Their limitation is that most collect via API rather than the consumer interface, and Yext explicitly documents that these can differ. Tool output is an estimate of visibility under that tool’s prompt set and collection method — not a census of the answers your buyers receive.

Is GEO just SEO renamed?

No, though they overlap. Google requires standard index and snippet eligibility for AI Overview inclusion, so technical SEO remains a precondition. But roughly a third of AI Overview citations come from outside the organic top ten, and ChatGPT, Perplexity, and Claude operate separate crawlers and retrieval systems from Google entirely.

Why measure mentions and citations separately?

Because they diverge. A brand recommended in the answer text without its domain in the source list has a different problem — and a different fix — from a brand cited as a source but never recommended.

What if the answers about my brand are simply wrong?

That is the most urgent finding a baseline can produce and the one most dashboards score as a positive, because they count the mention without evaluating it. Correcting misrepresentation precedes any attempt to increase visibility.

Run this on your brand.#

DeviLabs runs the full protocol as part of its fixed-scope GEO engagement: 296 captures, four engines, a 30-day measurement window, complete screenshot evidence, and a baseline report that states its own uncertainty.

You receive the raw capture workbook, not just a score. See the conversion half of that engagement in our Post-Peak case study.

Book a baseline call Download the template
References
  1. Aggarwal, P. et al. GEO: Generative Engine Optimization. arXiv, 16 November 2023 (KDD 2024)
  2. IAB. Measuring Visibility in the AI Era. 3 August 2026
  3. Kumar, P. et al. Generative Engine Optimization at Scale: Measuring Brand Visibility. arXiv, 18 June 2026
  4. A Measurement Framework for Generative Engine Optimization. arXiv, 29 April 2026
  5. Google Search Central. AI Features and Your Website. Updated 10 December 2025
  6. OpenAI. Overview of OpenAI Crawlers.
  7. OpenAI. Publishers and Developers FAQ.
  8. Perplexity. Perplexity Crawlers.
  9. Anthropic. Web Search Tool Documentation and crawler support article
  10. Microsoft. Manage Public Web Access for Microsoft 365 Copilot. 15 July 2026
  11. Yext. AI Citations, User Locations, and Query Context. 9 October 2025
  12. seoClarity. AI Overview Rankings Overlap. 17 July 2025
  13. Ahrefs. AI Overviews Reduce Clicks. 4 February 2026
  14. Pew Research Center. Google Users Less Likely to Click Links When AI Summaries Appear. July 2025
  15. Seer Interactive. Impact of AI Overviews on CTR. 21 May 2026
  16. Semrush. AI Overviews Study. 2025–2026