Generative Engine Optimisation · Buyer’s guide
How to choose a GEO agency in 2026
The answer
Judge a GEO agency on three things: whether they measure a baseline of your AI visibility before promising anything, whether the scope is weighted toward earned third-party authority rather than on-site tweaks, and whether they can show evidence for every deliverable they charge for. Walk away from guaranteed placements, week-two dashboards sold as outcomes, and retainers headlined by schema markup. Full disclosure: we sell GEO — so this guide is written to be used against us too.
Let’s handle the conflict of interest in the first paragraph: DeviLabs is a GEO agency. A guide to hiring one, written by one, deserves your suspicion.
So here’s the deal this page offers. Everything below is a test you can run on any vendor, including us — the questions have verifiable answers, the red flags reference published evidence, and nothing in here quietly concludes “therefore hire DeviLabs.” I wrote it because I keep watching founders sign bad retainers for reasons that take ten minutes to check, and because an informed buyer is genuinely easier to work with than a hopeful one.
Why this market is so noisy right now
GEO went from a research paper to a service category in roughly eighteen months. When a category grows that fast, supply gets created by relabeling: SEO agencies renamed their existing deliverables, dashboard startups positioned monitoring as strategy, and a wave of new shops formed around slide decks rather than method.
None of that makes the underlying discipline fake — the shift it responds to is measurable and large. Around 68% of US Google searches now end without a click, AI Overviews cut click-through by roughly 60% where they appear, and the sources AI engines cite overlap with Google’s top rankings by less than 20%. The market noise and the real opportunity are both loud; your job as a buyer is telling them apart.
The good news: bad GEO retainers are surprisingly easy to spot, because they share the same tell. They sell what’s easy to invoice instead of what the evidence says moves citations. Everything below is built on that one distinction.
The tell is always the same: deliverables optimised for the invoice, not the citation.
Six red flags, in the order you’ll meet them
Guaranteed placements in AI answers
Nobody controls a probabilistic model’s output. An engine’s answer to the same prompt changes between runs, between weeks, between model versions. A guarantee here is either ignorance of how the systems work or confidence you won’t check. Both disqualify.
Results promised in weeks
The work that moves recommendations — reviews, community presence, editorial mentions — accumulates on other people’s websites, on other people’s schedules. Three to six months to consistent movement is the honest timeline. “Visible results in 30 days” describes a dashboard being set up, not visibility being earned.
Schema markup or llms.txt as headline deliverables
Both are fine as hygiene. Neither moves citations: a 1,885-page study of schema additions found near-zero uplift, and major engines reportedly don’t even fetch llms.txt. When the cheap, easy items headline the scope, it’s because the scope was built backwards from what’s easy to deliver.
The dashboard is the product
Monitoring matters — we build measurement into everything. But a dashboard shows you the scoreboard; it doesn’t play the match. If the pitch spends more time on charts than on who’s going to earn you mentions in the publications engines cite, you’re buying reporting with a retainer attached.
No baseline before the promises
An agency that quotes you a strategy before measuring where you currently stand — which prompts name you, which name competitors, which sources the engines lean on in your category — is guessing. The baseline is a day of work. Skipping it means every later claim of progress is unfalsifiable.
A scope that never leaves your website
Roughly 83% of AI citations come from earned media — third-party coverage. A GEO scope that’s entirely on-site content and technical tweaks ignores the majority of the mechanism. It’s the single most common failure, because off-site work is the part agencies can’t fully control and hate to price.
Seven questions that expose everything
Ask these in the first call, in this order. You’re not testing whether the agency is good yet — you’re testing whether they’ll say true things under mild pressure.
- 01“How will you measure my baseline, and can I see the method?” — You want prompts, engines, cadence, and a written procedure. “We use a proprietary tool” without a method behind it is a weather app, not a measurement.
- 02“What share of the scope happens off my website?” — The honest answer is “most of it, eventually.” Listen for reviews, communities, digital PR. If the answer is all content and markup, see red flag 06.
- 03“What’s the evidence that each deliverable moves citations?” — The good ones cite studies and their own logged data. The rest cite “best practices.” Best practices in an eighteen-month-old field are folklore with a keynote slot.
- 04“What won’t work for us?” — Every honest practitioner has a list: categories where engines barely recommend brands, timelines that don’t fit a launch, budgets too small for the PR motion. No list, no trust.
- 05“When will nothing be happening, and what do I see during that?” — GEO has quiet months by design. You want interim proof: baseline by week two, content shipping, placements logged as they land.
- 06“Which engines matter for my buyers, and how do they differ?” — Engines share roughly 11% of citation domains. Anyone who talks about “AI” as one channel hasn’t measured more than one.
- 07“Do you run this on your own brand?” — The method is public or it isn’t. An agency selling AI visibility while invisible in AI answers about their own category is telling you how much they believe the pitch.
Question 03 is the one that does the most damage in practice. We published our own attempt at answering it — 24 common GEO claims checked against primary sources, fifteen of them false — as the GEO Claims Audit. Take it into your vendor calls; it was partly written for that.
What a fair scope actually contains
Strip away the packaging and a serious GEO engagement has four workstreams, in a predictable sequence and weighting.
| Workstream | What it looks like when real |
|---|---|
| Measurement (ongoing) | A documented prompt set across four engines, run on schedule, reported as share of voice against named competitors — not screenshots of one good answer. |
| Technical access (weeks 1–3, then done) | Crawler access verified, rendering fixed, entity details consistent. Finishes and falls off the invoice. If it’s still a line item in month six, ask why. |
| Answer-shaped content (monthly) | Pages built from the baseline’s real buyer questions, direct answers up front, specifics a model can quote. Volume matters less than fit to actual prompts. |
| Third-party authority (the long middle) | Pitches to the publications your engines already cite, review velocity programmes, genuine community presence. The slowest stream and the one that decides the outcome. |
The weighting is the fingerprint. Early months lean technical and content; by month four the majority of effort should sit in the third and fourth rows. A scope that stays 80% on-site forever is an SEO retainer wearing a lanyard. If you want the full mechanics of each stream, the playbook is public: how to get your brand recommended by ChatGPT.
What it should cost, and why
Honest ranges, EU/US market, 2026. Serious retainers for small and mid-size brands start in the low four figures per month. The cost driver isn’t software — it’s hours: measurement runs, content production, and above all outreach, which is human work that doesn’t compress.
Both extremes should worry you. A €300/month “AI visibility package” is a dashboard subscription with a logo on it — recall which 83% of the mechanism dashboards don’t touch. And a five-figure retainer is only justified by proportionally more earned-media output: more placements pitched, more reviews generated, more engines measured. More PDF is not more GEO.
One structural tip: ask for the technical phase priced as a project and the ongoing work as the retainer. It keeps the one-time work from silently annuitising — and an agency’s reaction to that request tells you plenty. Ours is on the pricing page.
How to hold any agency accountable — including us
Whoever you hire, run the engagement on falsifiable rails. Demand the baseline in writing by week two, with the prompt set attached so you can re-run it yourself. Get share-of-voice reporting against named competitors, per engine, on a fixed cadence. Have placements and mentions logged as they land, with links. And agree upfront what six months of no movement triggers — a strategy revision, a scope shift, or an exit.
An agency that resists any of that is planning to sell you effort instead of evidence. We’ve published our measurement method in full — the GEO Measurement Protocol — precisely so clients can check our work against it, and the free AI visibility audit that starts our engagements is the week-two baseline, delivered before you’ve paid anything. You keep it either way, including if you take it to a competitor. That’s not generosity; it’s the standard this post just told you to demand.
FAQ
What should a GEO agency actually deliver?
A measured baseline of your visibility across multiple AI engines before any promises; technical access fixes in the first weeks; answer-shaped content built around real buyer questions; sustained earned-media and review work on the third-party surfaces engines cite; and scheduled re-measurement with share-of-voice reporting. If the scope is mostly on-site tweaks, it’s a rebranded SEO retainer.
What are the biggest red flags when hiring a GEO agency?
Guaranteed placements in AI answers, results promised in weeks, schema markup or llms.txt as headline deliverables, a dashboard presented as the outcome, no baseline measurement before the pitch, and scopes that never mention third-party surfaces — the place roughly 83% of AI citations come from.
How much does GEO cost in 2026?
Serious retainers for small and mid-size brands typically run from the low four figures per month, because the work is labour-intensive: prompt-set measurement, content production, and digital PR outreach. Be suspicious of both extremes — a €300 “AI visibility package” is a dashboard subscription, and a five-figure retainer should come with proportionally more earned-media output, not a longer PDF.
Can an agency guarantee my brand appears in ChatGPT?
No. AI answers are probabilistic and nobody controls a model’s output. An honest agency commits to the inputs — measurement, content, third-party authority — and reports transparently on the visibility that follows. Anyone guaranteeing placement is guaranteeing something they don’t control.
Should my existing SEO agency just handle GEO?
Sometimes. The foundations transfer, and a good SEO team can learn the rest. Test them the same way you’d test a specialist: ask for a baseline measurement plan, ask what share of the scope lives off your website, and ask which claims they can support with evidence. If the answer is a keyword report with an AI logo on it, keep looking.
How long before a GEO engagement should show results?
Technical fixes land in the first weeks. Measurable share-of-voice movement typically takes three to six months, because third-party citations have to accumulate. A fair engagement defines interim proof: baseline delivered by week two, content shipping monthly, placements and mentions logged as they land, re-measurement on schedule.