Bar lengths are proportional to the count: Supported 4, Unsupported 5, False 15, of 24 claims audited.
How we ruled. A claim is Supported only where a controlled study or a primary source from the engine operator establishes it. Unsupported means the claim may be true but no controlled evidence exists — usually because the available data is observational and confounded. False means primary sources or controlled studies contradict it.
Every verdict below carries the study, the sample size, and the date. Where the evidence is a vendor study rather than independent research, we say so.
Most GEO advice is repetition. A claim appears in one vendor blog post, gets quoted in the next, and within a year it is treated as settled — without anyone checking the original source.
Some of these claims are true. Most are not, and a few are contradicted by controlled studies that were published and then largely ignored because the finding was commercially inconvenient.
We publish this because we sell GEO services, and the fastest way to be worth paying is to be the party that tells you which levers actually move.
The audit is versioned. When new evidence lands, verdicts change and the changelog records it.
The most repeated claim in GEO, and the one with the strongest evidence against it. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026, matched against control pages that never added it, and ran four separate statistical tests including a matched difference-in-differences analysis. No platform showed a meaningful citation increase. Google AI Overviews showed a 4.6% decline for treated pages — the only statistically significant result in the study.
OBSERVED
NO SCHEMA
SCHEMA · 3×
CONTROLLED
CONTROL
TREATED
The correlation was real; the cause was the site, not the markup. Observed: pages carrying schema were cited about three times as often. Controlled against matched pages of similar quality: the two bars are level.
Where the myth comes from: an earlier Ahrefs analysis of 6 million URLs found AI-cited pages roughly three times more likely to carry JSON-LD. That correlation is real. It is also explained by the fact that sites which implement schema tend to be well-maintained sites doing everything else right.
What is actually true: schema aids entity disambiguation, rich results, and knowledge graph association. Those are real benefits. Citation frequency is not one of them.
Source: Ahrefs, May 2026, 1,885 treated pages against matched controls
−4.6%
Change in AI Overview citations after adding schema
Google's Gary Illyes confirmed in July 2025 that Google does not support llms.txt and has no plans to. John Mueller compared it to the discredited keywords meta tag, and separately noted that server logs show AI crawlers do not even check for the file. As of Q1 2026 no major AI provider — OpenAI, Google, Anthropic, Meta, or Mistral — has publicly committed to reading it in production systems.
Adoption sits around 10% of domains in a study of 300,000 sites. One crawler analysis of over 500 million AI bot visits across 90 days found 408 requests targeting llms.txt directly.
What is actually true: llms.txt has emerging use in agentic and developer-tool contexts. It has no documented effect on search or answer-engine citation.
Sources: Google statements July 2025; SE Ranking adoption study; crawler log analysis, 2026
CLAIM 03
Allowing GPTBot makes your site eligible for ChatGPT Search#
FALSE
OpenAI documents these as separate crawlers with separate purposes. OAI-SearchBot governs whether a site can be surfaced in ChatGPT search results. GPTBot governs content that may be used to improve OpenAI's foundation models. ChatGPT-User is a user-triggered fetcher and controls neither.
Blocking OAI-SearchBot removes a site from ChatGPT search answers regardless of GPTBot status. This is the most consequential technical error we find in audits, because it is invisible — the site appears correctly configured to anyone who has not read the documentation.
Source: OpenAI crawler documentation, current 2026
CLAIM 04
Google-Extended controls whether you appear in AI Overviews#
FALSE
Google-Extended is a control for certain Google AI training and grounding systems. Googlebot remains the relevant crawler for Google Search and its AI features. Google states that eligibility for AI Overviews and AI Mode requires only that a page be indexed and eligible to appear with a snippet, with no additional technical requirements.
Source: Google Search Central, updated December 2025
Google removed FAQ rich results from Search on 7 May 2026, eliminating the traditional benefit. Google has never confirmed that FAQPage schema influences AI Overview or AI Mode citation, and the Ahrefs controlled study pooled FAQ schema with other types and found no citation effect overall.
Claims attributing specific percentage lifts to FAQ schema trace back to vendor experiments that changed page structure, content, and internal linking simultaneously.
Sources: Google Search, May 2026; Ahrefs controlled study, May 2026
CLAIM 06
Site speed and Core Web Vitals affect AI citation#
UNSUPPORTED
No controlled study isolates page performance as a variable in AI citation selection. Performance affects Google indexing and page experience, which is a precondition for AI Overview eligibility — so an indirect path exists. No published evidence establishes a direct effect.
Adding statistics, quotations, and citations improves visibility#
SUPPORTED
The foundational GEO paper — Aggarwal et al., arXiv, 16 November 2023, later accepted at KDD 2024 — tested nine content interventions against a benchmark of roughly 10,000 queries. Adding citations, quotations, and statistics, and improving fluency, produced the strongest visibility gains, reported at up to approximately 40% under test conditions.
The caveat that gets dropped: that figure is a relative, experiment-specific result on controlled rewrites, not a forecast for a commercial website. The paper reports that effectiveness varies by domain.
Source: Aggarwal et al., arXiv 2311.09735, ~10,000 query benchmark
Ahrefs analysed 1.4 million ChatGPT prompts and found that pages covering 26–50% of ChatGPT's internally generated sub-queries were cited more often than pages covering 100% of them. Depth on a focused scope outperformed breadth.
The "ultimate guide" format that dominated traditional SEO appears to work against citation when query relevance is held constant.
Source: Ahrefs, 1.4 million ChatGPT prompts, 2026
CLAIM 09
Headings that directly match the query increase citation#
SUPPORTED
In the same 1.4 million prompt analysis, pages with headings closely matching the query were cited 41% of the time, against 29% for weak matches. Cited pages showed a median title-to-prompt cosine similarity of 0.602 against 0.484 for uncited pages — rising to 0.656 when measured against the sub-queries ChatGPT generates internally.
STRONG MATCH
41%
WEAK MATCH
29%
Citation rate by heading match: strong 41%, weak 29%.
Heading structure was the strongest on-page lever tested, ahead of word count, topical breadth, and body copy.
Freshness is plausible and is widely assumed. No controlled study isolates it from the content changes that usually accompany an update. A page that is republished is typically also rewritten, restructured, and re-linked.
The original GEO experiment found information-rich rewrites outperformed keyword-oriented ones. Nothing in the subsequent literature supports repetition as a citation strategy, and Ahrefs found semantic similarity — not term frequency — predicts selection.
E-E-A-T is not a ranking factor in Google Search. Google has stated repeatedly that there is no E-E-A-T score, that the concept comes from the Search Quality Rater Guidelines, and that rater assessments do not directly affect rankings. Raters evaluate results to test and refine ranking systems.
Extending an already-misunderstood Search concept to AI answer engines, where no operator has published anything comparable, compounds the error.
What is actually true: the underlying qualities — accuracy, transparent authorship, genuine expertise — correlate with things that do get measured. The acronym is not itself a lever.
Source: Google Search Central and repeated Google statements
In controlled testing of ChatGPT citation signals, domain authority and backlinks showed no positive correlation with citation, and were slightly inversely correlated. ChatGPT appeared to evaluate content on relevance and structure rather than on authority signals.
This is the finding most at odds with how the SEO industry has approached GEO.
Overlap studies consistently show partial correlation at best. seoClarity reported approximately 32% of AI Overview citations overlapping Google's top ten. Ahrefs reported approximately 37.9% across 863,000 keywords and roughly 4 million URLs.
For ChatGPT the gap is far wider: only 6.82% of ChatGPT search results appeared in Google's top ten for the same queries, and only 16.61% appeared anywhere in Google's organic results.
OVERLAP WITH GOOGLE'S TOP TEN
AI OVERVIEWS · seoClarity
~32%
AI OVERVIEWS · Ahrefs
37.9%
CHATGPT · Ahrefs
6.82%
Bars are proportional: AI Overviews ~32% and 37.9%; ChatGPT 6.82%.
What is actually true: ranking correlates with citation and is worth having. It is neither necessary nor sufficient, and for ChatGPT specifically it explains very little.
This is the most misreported finding in AI search. The headline statistic — Reddit as the most-cited domain — is true for some engines and flatly false for others.
Analysis of OpenAI's search behaviour found Reddit pages appearing as candidate sources in 76% of ChatGPT searches, but only 0.61% were selected — 491,024 Reddit pages retrieved, 3,012 cited. A 99.39% rejection rate. For Claude, one analysis found zero Reddit citations across 139,601 grounding sources between May and July 2026; Reddit was not supplied as a candidate at all.
Gemini, Google AI Mode, Google AI Overviews, and Perplexity do cite Reddit heavily. Those four run live web retrieval.
Gemini
CITES REDDIT
AI Mode
CITES REDDIT
AI Overviews
CITES REDDIT
Perplexity
CITES REDDIT
ChatGPT
DOES NOT
Claude
DOES NOT
CHATGPT RETRIEVAL FUNNEL
RETRIEVED
491,024
CITED
3,012
ChatGPT retrieved 491,024 Reddit pages and cited 3,012 of them — 0.61% selected, 99.39% rejected. Four engines cite Reddit (Gemini, AI Mode, AI Overviews, Perplexity); two effectively do not (ChatGPT, Claude).
What is actually true: a Reddit strategy targets four engines and is ignored by two, including the two that favour editorial and primary sources. Anyone quoting a single cross-engine Reddit percentage is averaging across systems that behave in opposite ways.
Wikipedia is among the most-cited domains, particularly for ChatGPT. But citation share moved sharply within weeks: ChatGPT's Wikipedia citation rate fell from roughly 55% of prompt responses to under 20% in mid-September, while remaining stable on other engines.
Being cited from Wikipedia is also not the same as a brand benefiting. The citation credits Wikipedia.
The direction is consistent across many datasets. Semrush reported 4.4x. Ahrefs found AI traffic at 0.5% of sessions driving 12.1% of signups. Shopify reported roughly 50% higher conversion on AI-referred sessions. Visibility Labs found a far smaller 1.3x across 94 ecommerce brands. BrightEdge found the opposite for Fortune 100 brands through August 2025.
Why it is still unsupported: every one of these is observational. The conversion event is defined differently in each — form fills, signups, purchases — so the multiples are not comparable. And the causal story runs backwards: AI users arrive later in their decision process, which the AI did not cause. Selection effect, not treatment effect.
What is actually true: AI-referred visitors behave differently and are worth tracking separately. No controlled study establishes that AI referral causes higher conversion.
Sources: Semrush 2026; Ahrefs June 2025; Shopify Q1 2026; Visibility Labs 2025; BrightEdge September 2025
Pew Research Center found users clicked a traditional result in 8% of visits where an AI Overview appeared against 15% where none did, across 68,879 searches from 900 US adults. Ahrefs found AI Overview presence associated with 58% lower average CTR for the top-ranking page. Seer Interactive reported CTR falling from 1.76% to 0.61% across its tracked queries.
Three independent datasets, three methodologies, same direction. This is the best-evidenced claim in the audit.
Sources: Pew Research Center, July 2025; Ahrefs, February 2026; Seer Interactive, 2026
Semrush tracked AI Overviews on 6.49% of its keywords in January 2025, 24.61% in July 2025, and 15.69% in November 2025. Published figures elsewhere range from roughly 15% to 60%.
These are not contradictory. They measure different keyword panels, countries, devices, dates, and query mixes. Any single percentage presented without those parameters is not a measurement.
The number is real and comes from Aggarwal et al. It is the maximum relative improvement observed for the strongest intervention, on controlled rewrites, against a benchmark query set, in 2023. It is not an expected outcome for a commercial website, and the paper reports variation by domain.
Quoting it as a service benchmark misrepresents the source.
ChatGPT cited Reddit in close to 60% of prompt responses in early August before falling to around 10% by mid-September. Wikipedia fell from roughly 55% to under 20% in the same window on ChatGPT while holding steady on other engines.
CHATGPT REDDIT CITATION SHARE · 6 WEEKS
Reddit's share of ChatGPT prompt responses falls from roughly 60% in early August to roughly 10% by mid-September.
Separately, a 2026 study found identical prompts producing distinct outputs roughly 25% of the time on GPT-4o-mini, and query wording accounting for approximately 26.5% of response variance.
What is actually true: citation position is a rate with an error range, measured over repeated runs. A single observation is not a state.
Yext explicitly documents that citations returned by APIs may differ from those shown in the consumer application. Most commercial tools collect via API because it is easier to automate.
The divergence between engines is also larger than most reporting suggests: one cross-platform analysis found only 11% of cited domains overlapping between ChatGPT and Perplexity for identical queries.
What is actually true: each tool reports visibility under its own prompt set, collection method, product mode, and citation definition. Two tools disagreeing does not mean one is wrong.
Sources: Yext, October 2025; Pixis cross-platform analysis, 2026
Ahrefs found that roughly half of URLs ChatGPT retrieves are cited, and in a separate cut, that 85% of retrieved pages went uncited. Retrieval and selection are distinct stages with distinct determinants.
RETRIEVAL → CITATION
RETRIEVED
100%
CITED
~50%
Two stages: of the pages ChatGPT retrieves, roughly half are cited — the bar drops by about half between retrieval and citation.
A site can be perfectly crawlable, correctly indexed, retrieved for the right query, and still never appear in the answer.
ChatGPT generates internal sub-queries — fan-out — when answering. 95% of those fan-out queries had zero monthly search volume by conventional metrics, and 32.9% of cited pages appeared only in results for a fan-out query rather than the original prompt.
Keyword tools built on search-volume data are structurally blind to roughly a third of citation opportunities.
Four claims out of 24 are supported by controlled evidence:
01
Adding statistics, quotations, and source citations improves visibility.The foundational GEO paper, ~10,000 query benchmark. Domain-dependent, and the 40% figure is experiment-specific.
02
Headings that closely match the query increase citation rate.41% against 29%, across 1.4 million ChatGPT prompts. The strongest on-page lever tested.
03
Focused scope outperforms comprehensive coverage.Pages covering 26–50% of sub-queries were cited more than pages covering all of them.
04
AI Overviews reduce clicks to websites.Three independent datasets, three methodologies, consistent direction.
Notice what these have in common. None of them is technical. None involves markup, files, or configuration. All four are about whether the content answers a specific question clearly and is structured so the answer can be lifted out.
That is a less saleable conclusion than a schema audit. It is the one the evidence supports.
Is it a controlled study or an observation? Nearly every large GEO statistic is observational. Sites that added schema also improved content. AI-referred users were already further along. Correlation is the default state of this field.
02
What is the denominator? 32%, 37.9%, and 90% all describe Google overlap and none can be compared to the others. They count different things.
03
Which engine? Reddit is a primary source for four engines and filtered out entirely by two. Cross-engine averages hide opposite behaviours.
04
Who paid for it? Vendor studies are not worthless — several cited here are vendor studies and they are the best data available. They are also drawn from client bases and rarely replicated.
05
When? ChatGPT's Reddit citation share fell from 60% to 10% in roughly six weeks. A statistic from last quarter may describe a system that no longer exists.
We sell GEO services. That is a reason to read this sceptically, so here is what the audit means for what we sell.
It means we do not sell schema packages as a citation lever, because the controlled evidence says they are not one. It means we do not quote the 40% figure. It means our measurement work reports rates with error ranges rather than scores, because citation shares move by tens of percentage points within weeks.
It also means the work is narrower and less impressive-sounding than the category norm: make sure the engines can retrieve you, make sure what they say about you is factually correct, and structure content so a specific question gets a liftable answer.
Four levers, not forty. We would rather be right about four.
Not measurably, according to the only controlled study on the question. Ahrefs tracked 1,885 pages that added JSON-LD against matched controls and found no meaningful increase on any platform, with a small statistically significant decline on Google AI Overviews. Schema retains value for entity disambiguation and rich results — citation frequency is not among its documented effects.
Should I implement llms.txt?
There is no evidence it affects search or answer-engine citation. Google has stated it does not support the file and no major AI provider has committed to reading it in production. It costs almost nothing to publish, so the honest answer is that it is harmless and unproven rather than actively useful.
Do backlinks help AI visibility?
Controlled testing of ChatGPT citation signals found domain authority and backlinks showed no positive correlation, and were slightly inversely correlated. Backlinks remain valuable for traditional search. The evidence does not extend them to AI citation.
Does AI traffic really convert better?
Many datasets point that way, but all are observational and define conversion differently, and at least one large study found the opposite for enterprise brands. The likeliest explanation is that AI users arrive later in their decision process — a selection effect the AI did not cause.
Is being cited by Reddit good for my brand?
It depends entirely on which engine you care about. Gemini, Perplexity, Google AI Mode, and AI Overviews cite Reddit heavily. ChatGPT rejects over 99% of the Reddit pages it retrieves, and Claude does not appear to retrieve Reddit at all.
What actually works, then?
Four things have controlled evidence behind them: information-rich content with statistics and sources, headings that closely match the question being asked, focused scope over comprehensive coverage, and verified crawler access. Everything else in this audit is either unproven or contradicted.
DeviLabs runs a documented GEO baseline: 296 captures across four engines, a 30-day measurement window, and a report that states its own uncertainty. No scores without denominators, no claims without evidence.