GEO for Ecommerce: What Actually Gets You Cited
We audited the five guides ranking for this query against the controlled evidence. Here is what held up.
In this article
- What GEO for ecommerce actually means
- What the guides ranking for this query recommend
- The most-recommended tactic has a null result behind it
- What the controlled evidence actually supports
- Off-domain work: our position, not a tested finding
- What Shopify already does, and what it does not
- Product feeds and agentic checkout
- Presence is not accuracy, and DTC loses on accuracy
- How do you measure this without fooling yourself?
- Where the opening is, and what we cannot prove
- Frequently asked questions
- Sources
GEO for ecommerce is the work of getting AI engines to name your products when a shopper asks a buying question. Almost every guide answers that with structured data. The one controlled test of that question found no citation lift on pages engines were already citing. What actually survived testing is narrower and harder to sell: answer-shaped content and tightly scoped pages. Everything else here, including the off-domain work we think matters most, is argued rather than proven, and this article says which is which.
What GEO for ecommerce actually means
The term covers two jobs that behave differently and fail differently. One is being cited when a shopper asks a question your category answers. The other is being described correctly when an engine already knows you exist. Most DTC brands assume they have the first problem and turn out to have the second.
Generative Engine Optimization (GEO) is the practice of changing the evidence AI engines draw on about your brand, so the answers they generate name you and describe you correctly.
The commercial pressure is measurable. Pew Research Center tracked 68,879 Google searches from 900 US adults and found people clicked a traditional result in 8% of visits where an AI summary appeared, against 15% where none did.
What the guides ranking for this query recommend
We read the five guides ranking for "geo for ecommerce" on 24 September 2026 and scored each against five checks, counting a recommendation only where the guide tells the reader to do the thing. All five are linked in the sources, so you can check the grid. The agreement is high and the evidence under it is thin.
| Guide | Schema as a citation lever | Cites a controlled study | Metric with a denominator | Accuracy scored apart from presence | Volatility or error range |
|---|---|---|---|---|---|
| Hello Retail | Yes | Yes | Named, no formula | No | No |
| Salsify | No | No | No | No | No |
| Mirakl | Yes | No | Named, no formula | No | No |
| ShopVision | Yes | No | Named, no formula | No | No |
| Yotpo | Yes | No | Named, no formula | No | No |
Two things the grid does not hold against them. Three of the five send you off your own domain, to review platforms, marketplaces or Reddit, so the third-party idea is not missing here. And Salsify warns that outdated content loses the recommendation outright, which is closer to the accuracy problem than anyone else gets. None of them turns that warning into something measured. Read the middle column strictly: it asks for a published formula, not a metric named in passing.
The most-recommended tactic has a null result behind it
Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 and matched them against control pages that did not. That is a difference-in-differences design, built to isolate what the markup caused rather than to observe what happened to sit alongside it, and we have not found another test of this question built that way.
No platform showed a meaningful citation increase. AI Mode and ChatGPT drifted a couple of points upward, close enough to zero to be noise. Google AI Overviews moved in the other direction, and the authors were careful to say they could not confidently attribute even that to the markup.
There is a limitation the study states plainly and most people repeating it leave out. Every page in that dataset already had more than a hundred AI Overview citations before any schema was added. These were pages the engines had already decided to surface. The authors say directly that schema might still matter for pages not yet cited at all, which is where most DTC brands are standing. So the honest reading is narrower than the headline. The test is strong evidence against schema as a lever for a brand already inside the consideration set. It does not license a store with no citations at all to skip it.
Our own GEO Claims Audit records this as Claim 01 and marks it False. Not because schema is useless: it helps with entity disambiguation, rich results and knowledge graph association, all worth having. Citation frequency is not one of them.
The timing matters here. That test published in May 2026. Hello Retail's guide, the best-researched page ranking for this query, published a month later and still tells readers that "sites with structured data and FAQ blocks saw a 44% increase in AI search citations", citing a vendor report. The same page also notes that 65% of pages cited by Google AI Mode carry structured data. That second number is the more interesting one, and it is a co-occurrence on pages that got cited, not a lift caused by the markup. The controlled test is what separates those two readings, and it went the other way.
Schema work is the easiest thing in this category to buy: a deliverable, a before and after, a developer who can finish it in a sprint. The work that moves citations has none of those properties, which is why it is missing from most proposals.
What the controlled evidence actually supports
Our audit tested 24 claims against controlled studies and primary sources from the engine operators, and four came out supported. Three of them are things you can do to a page. The fourth is a fact about the market rather than a lever. Not one of the four is technical.
- Put the buying question in the heading, the title and the URL. Headings that closely match the query are the strongest on-page lever the audit tested. The same pattern holds one level up: the Ahrefs analysis of 1.4 million ChatGPT prompts found pages whose titles and URLs align semantically with the model's internally generated fan-out queries are the ones that get cited. Vague category headings align with nothing.
- Keep the page narrow. Our claims audit reads the same Ahrefs dataset as showing that pages covering part of a query's fan-out were cited more than pages covering all of it. We have not found that breakdown published on the study page itself, so treat it as the audit's reading rather than a figure you can quote.
- Cite real sources and real numbers. This is the on-page lever with a controlled benchmark behind it, the Princeton GEO paper. We do not quote its headline percentage, because it measures controlled rewrites in a 2023 experiment and it is not a forecast for a commercial store.
- The fourth is not an on-page lever. It is the finding that AI Overviews reduce clicks to websites, the best-evidenced claim in the audit and the reason the three above it are worth the effort.
Retrieval is not citation, and the gap between them is where most ecommerce content disappears. In that same 1.4M-prompt study, Reddit content was cited just 1.93% of the time while making up 67.8% of the pages the model pulled in and never named. A dashboard counting mentions cannot tell you which of those happened to you.
Off-domain work: our position, not a tested finding
Notice what is missing from that list. The thing we tell clients matters most, being talked about somewhere you do not own, is not one of the four. We are going to be exact about that rather than quietly promote it, because the article would otherwise be doing what it criticises.
Our own audit marks the backlinks-and-domain-authority claim False, and marks the Reddit claim False as stated, because the effect is engine-dependent: several engines cite Reddit heavily while ChatGPT rejects almost all of what it retrieves from there. So off-domain corroboration is a position we hold from running this work, and it is not something a controlled study has established for you.
The reasoning: your product page is a claim, a third-party page is a witness. An engine weighing two similar stores has nothing to separate them except what people who work for neither have written. In practice that is three standing jobs: a post-delivery email asking fulfilled customers for a review on the platform your category is actually judged on rather than on your own site; a monthly pitch to the writers of the comparison pages already ranking for your category, offering test units or data rather than an affiliate rate; and a named, disclosed account answering questions in the two or three forums your buyers use. You cannot write any of these yourself, which is the point. Faking them gets detected and poisons the signal you were building.
What Shopify already does, and what it does not
Before paying anyone to add Product schema, check what your theme already emits. Shopify themes generate product structured data through a built-in Liquid filter, and the free Dawn theme uses it, so a standard store is usually already carrying most of what an agency will quote you for.
The documented output covers name, description, image, brand, category, URL and an offers block with price, currency and availability. GTIN, SKU and review or rating data are not in it. Those three are the only fields worth paying someone to add.
That gives you a concrete test to run this afternoon. Fetch one product URL the way a crawler would, without a browser, and read the raw HTML that comes back. If price, availability and variant names are in that response, an engine can read them. If they only appear after JavaScript runs, several AI crawlers execute little or no JavaScript and will see an empty shell. App-injected review widgets are the usual offender: the stars render for a shopper and are absent from the source.
The other Shopify-specific check is access. Shopify serves a default robots.txt you can override with a robots.txt.liquid template, and the crawlers you care about are not one list. OpenAI documents four separately, of which two decide whether ChatGPT can see you: GPTBot handles training, and OAI-SearchBot is the one that governs search answers. Anthropic splits its crawlers the same way, and this is where the mistake repeats: ClaudeBot is the training crawler, while Claude-SearchBot is the one that governs search answers. PerplexityBot covers Perplexity.
Google is the exception worth understanding, because the obvious answer is wrong. Google-Extended is a robots.txt token governing whether your content trains and grounds Gemini, and Google's crawler documentation states it has no effect on Google Search. AI Overviews and AI Mode are Search surfaces, gated by ordinary Googlebot access. Blocking Google-Extended costs you Gemini, not AI Overviews.
Run curl -s https://yourstore.com/robots.txt and look for a Disallow under any of those agents. The most common technical error we find in audits is a site allowing GPTBot and assuming it is therefore eligible for ChatGPT Search, when the crawler that decides that is OAI-SearchBot. Plenty of stores also blanket-blocked AI crawlers in 2023 and 2024, when the argument was about training data, and never went back.
Product feeds and agentic checkout
This is the genuinely ecommerce-specific layer, and it is newer than the evidence base. OpenAI now publishes an Agentic Commerce Protocol with a product feed specification, so merchants can submit structured catalogue data rather than waiting for a crawler to infer it. Google runs an equivalent path through Merchant Center.
Nobody has controlled evidence that being in these feeds changes how often an engine recommends you, because the programmes are too new for a matched test to exist. What a feed does do is remove the parsing step from price, stock and variant data, which is the category of fact engines get wrong most often. So we treat feeds as accuracy plumbing, not a fifth lever, and we would not let a vendor price them as a visibility product until someone publishes a test.
Presence is not accuracy, and DTC loses on accuracy
Every guide in the grid treats a mention as a win. The failure that costs a store money is the engine that names you and then gets you wrong: the discontinued variant, last year's price, a shipping policy you changed, or a similarly named company in another country.
Presence was never the issue in those runs. Accuracy was. That is why we score factual accuracy separately from presence, and why correction comes before amplification in every engagement. Every prompt is scored against five binary criteria fixed before measurement: correct brand name, correct location, correct core services, correct people, and not confused with another entity. That last one is the criterion no public framework includes, and in practice it fails more often than the other four combined. A mention-rate tool scores every one of those as a success.
How do you measure this without fooling yourself?
Define the unit before you measure it. Every number in an AI visibility report is a ratio, and a ratio without a stated denominator cannot support a decision. Our measurement protocol publishes the definitions, and the standing baseline is 296 captures split between identity prompts and commercial prompts, kept apart because they are different problems.
| Metric | Definition |
|---|---|
| Mention rate | Runs where the brand name appears ÷ valid runs |
| Citation rate | Runs where the brand domain appears in sources ÷ valid runs |
| Share of voice | Brand mentions ÷ all brand mentions in the same run set |
| Accuracy rate | Runs with factually correct brand claims ÷ runs mentioning the brand |
| Volatility | Repeat runs producing a different outcome for the same prompt |
Volatility is the one to insist on, and it is reported alongside every rate in our protocol. A 40% mention rate with 5% volatility and a 40% mention rate with 30% volatility are not the same finding, and only one of them supports a decision. A single run of a single prompt tells you almost nothing.
For traffic, build a GA4 channel group matching the AI hostnames and treat the result as a floor. It will never catch Google AI Overviews, whose clicks arrive with a plain google.com referrer and land in organic search. You will also meet the claim, repeated across this category, that AI-referred traffic converts better than organic. The datasets behind it are observational, each defines conversion differently, and the causal story runs backwards: AI users arrive later in their decision process, which the AI did not cause. Selection effect, not treatment effect. Track the segment separately and measure your own store.
Where the opening is, and what we cannot prove
AI recommendation is not winner-take-all the way Google's first page is. Engines evaluate authority per query, so a small brand with a genuinely specific answer can be named ahead of a retailer many times its size. That opening is real and temporary, and it closes as the category gets crowded.
Two honest caveats before you act on any of this. The three-to-six-month timeline in the FAQ is our read from the engagements we run, not a finding from a controlled study, and it is marked that way on purpose. And DeviLabs has not published a GEO case study, because the practice is new and we do not publish results we cannot measure. What sits behind this article is the external research cited below and a published protocol, not a client win we are asking you to take on trust.
Start with the cheap half this week. Check whether each engine's search crawler can actually reach your product pages, fetch one PDP without a browser and confirm the price is in the response, then ask ten real buyer questions across four engines and write down every fact they get wrong about you. That list is your first month of work, and it is where our own GEO engagements begin too.
Frequently asked questions
Keep it, for rich results, entity disambiguation and knowledge graph association. Those benefits have evidence behind them. What the controlled testing does not support, on pages an engine already cites, is buying schema work as a way to increase AI citations, which is how most GEO packages are sold.
The evidence is the same; the surfaces are not. A store has product pages, variants, prices, shipping rules and review profiles, and every one is a fact an engine can state wrongly. It also has feeds, which no service business has to think about.
Our read, from our own engagements rather than a controlled study: access fixes land in days, content starts getting retrieved in weeks, and the third-party layer shows consistent movement at three to six months. Treat any agency promising faster than that with suspicion.
Not to start. Write ten real buyer questions, run them in ChatGPT, Perplexity, Google and Claude from fresh logged-out sessions, and log what comes back each week. A tool saves time later. It will not tell you which prompts matter in your category.
It removes you from live retrieval, which is separate from training. OpenAI documents GPTBot and OAI-SearchBot as different crawlers with different purposes, so blocking the search crawler takes you out of ChatGPT search answers whatever you do with the other one.
No. The GEO practice is new and we do not publish results we cannot measure, so there is no client case study behind this article. What sits behind it is the published measurement protocol, the claims audit, and the external studies cited in the sources below.
Sources
- We tracked 1,885 pages adding schema. AI citations didn't move — Ahrefs, May 11, 2026. Cited for: “1,885”
- Why ChatGPT cites one page over another (study of 1.4M prompts) — Ahrefs, Apr 15, 2026. Cited for: “1.93%”
- Google users are less likely to click on links when an AI summary appears in the results — Pew Research Center, Jul 22, 2025. Cited for: “8%”
- GEO for ecommerce: how to get your products cited by AI search engines — Hello Retail, Jun 15, 2026. Cited for: “44%”
- GEO Claims Audit — DeviLabs, Aug 11, 2026. Cited for: “24 claims”
- Liquid filters: structured_data — Shopify. Cited for: “structured_data”
- Agentic Commerce Protocol — OpenAI. Cited for: “Agentic Commerce Protocol”
- OpenAI crawlers and user agents — OpenAI. Cited for: “OAI-SearchBot”
- Google crawlers and fetchers — Google. Cited for: “Google-Extended”
- Does Anthropic crawl data from the web, and how can site owners block the crawler? — Anthropic. Cited for: “Claude-SearchBot”
- GEO Measurement Protocol — DeviLabs, Aug 8, 2026. Cited for: “296 captures”
- SEO, GEO, AEO: generative engine optimization for ecommerce (scored in the grid above) — Salsify. Cited for: “GEO”
- What is eCommerce GEO? (scored in the grid above) — Mirakl. Cited for: “GEO”
- Generative engine optimization for ecommerce: the 2026 GEO playbook (scored in the grid above) — ShopVision. Cited for: “GEO”
- What is GEO? The complete guide (scored in the grid above) — Yotpo. Cited for: “GEO”
How this article was made: researched and drafted with AI assistance, built on DeviLabs’ published methods and Dominykas Jankauskas’s own positions, then fact-checked against the primary sources listed above and reviewed before publication. Read the full editorial policy. Spotted an error? Email info@devilab.eu.
Find out what AI says about you
Free 30-minute consultation call. We go through how AI engines pick the brands they name in your category, where the opening is for you, and what it would take to get into those answers.