The GEO Playbook for Agencies
How an agency sells GEO without overpromising: the method, the deliverable, the pricing, and the four client objections you will meet.
When an agency sells Generative Engine Optimization, it is selling three things a client cannot get alone: a measurement they can trust, a diagnosis of why they are absent from AI answers, and the editorial work to fix it. It is not selling a ranking, because there is no ranking. That difference decides your method, your deliverable and your price.
This is the playbook we would hand a strategist on day one. It is built on the studies that exist rather than the ones people quote, which means part of it is about what you should refuse to promise.
The short version
- You are selling measurement and editorial work, not a lever. No technique has yet shown a causal, stable, cross-platform effect, so a promise of “position 1 in ChatGPT” is a promise you cannot keep.
- The demand is real and it is B2B. 51% of software buyers now start a purchase in an AI chatbot, and 64% meet inaccurate recommendations often (G2, 2026, 1,076 buyers).
- Accuracy sells better than visibility. “Your competitors are being recommended” is a maybe. “The AI is quoting a price you retired” is a fire.
- Your method is the product. A client can buy a $29 tool. What they cannot buy is a frozen prompt set, a control page and someone who reads the answers.
- Retainer, not project. AI answers drift month to month, so a one-off audit measures noise. Sell the baseline once and the tracking monthly.
Can you actually charge for this?
Yes, and the market has already priced the tooling layer, which tells you where your margin is.
| Layer | What it costs | Who captures it |
|---|---|---|
| Monitoring tool | $29 to $300 a month, verified August 2026 | The vendor |
| Prompt design and baseline | 8 to 20 hours of senior time | You |
| Diagnosis and editorial plan | Recurring | You |
| Content and earned mentions | Recurring | You |
| Reporting a client can defend upward | Recurring | You |
The tool is the cheapest part of a GEO engagement and the easiest to replace. If your proposal is a dashboard resold at a markup, you are selling the one component with no moat. The defensible product is the prompt set, the interpretation and the editorial work.
A useful way to price: your fee should be a multiple of the tool cost that reflects the senior hours in the diagnosis, not a percentage of media spend. GEO has no media spend, which is exactly why the percentage-of-spend model that funds paid teams does not transfer.
The method, in five parts
This is the deliverable. Give it a name, run it the same way every time, and let the repeatability be what the client is buying.
1. Freeze a prompt set
Between 20 and 60 prompts, written the way a buyer speaks, not the way a keyword tool writes. Four families:
- Category discovery. “What is the best [category] for [segment]?”
- Brand direct. “What does [brand] do?” and “Is [brand] any good?”
- Comparison. “[Brand] vs [competitor]”, which is where the sharpest damage shows up.
- Objection. “Is [brand] expensive?”, “Does [brand] work for [use case]?”
Freeze them. The moment you improve a prompt, you lose the ability to compare with last month, and comparison is the entire job.
2. Record the conditions, not just the answer
Engine, model version, country and language. The same prompt answers differently on all four, and a report that omits them cannot be reproduced by the client or by you in three months.
Language matters more than most teams expect. In Temso’s analysis of over 7 million citations, French prompts were answered mostly with French-language sources: 82.1% in AI Overviews, 87.2% in Copilot, 72.3% in ChatGPT. If your client sells in a market whose language they do not publish in, that is your first finding, and it is usually the biggest one.
3. Run each prompt several times
AI answers vary between identical runs. One check per prompt per month produces a chart that is partly measuring randomness, and a client who spots that once will not trust the next report either.
We measured it: three runs of one question on AI Overview returned 11 cited domains, of which 4 appeared only once. Run each prompt several times and report the aggregate. This is the single practice that separates a defensible engagement from a plausible one, and it is the reason to use a tool rather than a person with a browser.
4. Read the sources, not only the mentions
Whether the brand appeared is the shallow layer. The useful layer is which pages the model leaned on. That list is your editorial brief: it tells you which third-party pages own the category, which of your client’s pages are already trusted, and which competitor is being quoted instead of them.
5. Change one thing, against a control
Pick a page, improve it, leave a comparable page untouched, and measure both against the frozen prompt set over several weeks.
Without a control you cannot separate your work from the platform’s own drift, and platform drift is large. This is the part clients rarely ask for and the part that makes your report defensible when a CFO asks what changed.
What to promise, and what to refuse
The evidence here is thinner than the industry admits, and knowing the line is a commercial advantage rather than a limitation. A July 2026 review of 45 GEO studies concluded that no technique has demonstrated a causal, stable, longitudinal, cross-platform effect.
Safe to promise:
- We will tell you what the major engines currently say about you, with the sources.
- We will tell you where the answer is wrong, and that is fixable.
- We will improve the pages the models are already reading, and measure against a control.
- You will be able to reproduce every number in this report.
Refuse, whatever the client asks:
- A position. There is no position to sell.
- A timeline for appearing. No study establishes one, and index refresh cycles are outside anyone’s control.
- A citation lift from schema markup. Ahrefs tested it against a control group, on 1,885 pages against roughly 4,000, and found no significant increase, with a 4.6% decrease in AI Overviews. Add structured data because it describes the page, not as a GEO lever. Our own sweep of the pages an AI actually cited found 93% carrying some JSON-LD and 41% declaring what the page is, which has no uncited comparison group and so can only say that heavy markup is not what those pages have in common.
- Anything based on “LLMs prefer 40 to 60 word blocks” or a specific ideal word count. No primary evidence exists for either.
What does have controlled support is narrower and more useful than the folklore. In the Princeton GEO-bench experiment over 10,000 queries, adding quotations raised a source’s share of the answer by 41%, statistics by 30% and cited sources by 27%, while keyword stuffing scored below doing nothing. That is an editorial brief you can hand to a writer today.
The four objections you will meet
“Isn’t this just SEO?” No, but it depends on SEO completely and you should say so first. Retrieval runs on a search index, so a page that cannot be crawled or ranks for nothing is not a candidate. GEO is a layer above. Selling it as a replacement is how agencies lose the account in month four.
“How do I know it worked?” Because you kept a control page and a frozen prompt set. If you cannot answer this in one sentence, you have not designed the engagement properly.
“Our traffic from AI is tiny.” It is, and that is the wrong number. AI referrals are still a small share of sessions, but Adobe Analytics, measuring over a trillion visits, found they convert 42% better. Meanwhile Pew found that when Google shows an AI summary, 8% of visits end in a click against 15% without. Fewer clicks, better clicks. Judging this channel by sessions will mislead the client in both directions.
“Can’t we just do this ourselves?” They can run the tool. What they will not do is freeze a prompt set for six months, resist improving it, keep a control page, and read the cited sources every month. Say that plainly. It is true, and it describes the work you are charging for.
What the report should contain
A GEO report that survives being forwarded to a CFO has six sections. Anything else is decoration.
| Section | What it answers |
|---|---|
| Mention rate | In what share of the frozen prompts does the brand appear at all |
| Share of voice | Against the three competitors the client actually names in sales calls |
| Position and framing | Named first, named last, or named as the cheap option |
| Accuracy log | Every wrong statement the engines made, with a screenshot and a date |
| Source map | Which pages the answers leaned on, ours and theirs |
| What we changed, and the control | One line per change, with before and after |
The accuracy log is the section clients forward internally, and it is usually what renews the retainer. A competitor beating you in an answer is an argument. A model telling buyers something false about your client is an incident, and incidents get budget.
Notes by vertical
The category behaves differently by sector, and these are the differences we see most often.
Higher education. The buyer is 17 or 18, the decision is worth tens of thousands, and the question is asked in the student’s own language. Combined with the language finding above, an institution recruiting internationally while publishing only in its own language is structurally invisible to a large part of its market. The accuracy risk is also unusually sharp: a model inventing a tuition fee or an entry requirement is a compliance problem, not a marketing one.
Travel. Answers go stale faster than anywhere else, because prices, routes and entry rules change under them. Recency work matters more here than in any other vertical, and the failure mode is not absence but confident obsolescence.
Professional services and consultancies. The queries are geographic and reputational: “best [discipline] consultancy in [city]”. Third-party lists and directories carry the category, so the work is more earned-mention than on-page.
Ecommerce and retail. Comparison prompts dominate and price accuracy is the whole game. A retired price quoted confidently by an assistant is the most common finding, and the easiest win to demonstrate in month one.
Common Questions From Agencies
What should an agency charge for GEO?
There is no published benchmark, and anyone quoting a standard rate is inventing it. What is verifiable is the input cost: monitoring tools run from $29 to $300 a month. Price your fee from the senior hours in prompt design, diagnosis and editorial work, not as a markup on the tool, because the tool is the replaceable part.
Should we resell a tool or build reporting on top of one?
Resell only if the tool lets you export raw data. The risk is concentration: this category already saw one vendor shut down in October 2025 and another acquired in June 2026. Ask where the historical data lives and whether you can take it with you before you build a client-facing report on it.
How long before a client sees results?
There is no established figure, and a study of 45 papers found no technique with a demonstrated stable effect. Set the expectation on the measurement instead: a defensible baseline takes a few weeks of repeated runs, and a change measured against a control needs several weeks after that.
Can GEO work be white-labelled?
The reporting can. The method should not be hidden, because the method is the reason the client believes the numbers. Agencies that publish how they measure tend to be the ones quoted by other agencies, which is its own acquisition channel.
Is GEO a separate service line or part of SEO?
Commercially it sells better as a separate line, because it has a different deliverable and a different buyer conversation. Operationally it cannot be separated from SEO, because retrieval depends on the index. Sell it separately, staff it together.
What if the client is already invisible everywhere?
That is the easiest engagement to prove, not the hardest. A brand at zero has nowhere to go but up, and the baseline is unambiguous. The difficult client is the one already appearing sometimes, because separating your work from normal variation there requires the control discipline above.
Do we need to track engines beyond ChatGPT?
Track the engines the client’s buyers use, which is measurable rather than a matter of taste. In English-speaking markets ChatGPT and Google AI Overviews cover most discovery. If the client sells in France, Spain or China, the answer changes, and engines like Doubao and Qwen become relevant in ways an English-market comparison will not surface.
Where this leaves you
The agencies that will own this category are not the ones with the best-looking dashboard. They are the ones who can answer “how do you know?” without changing the subject.
That means a frozen prompt set, several runs per number, a control page, and a written refusal to promise a position that does not exist. It is a less exciting pitch than the one your competitors are making this quarter, and it is the one that survives the second renewal.
Next: what GEO is and what the research supports, and the tool comparison with prices we verified.
Run a client's brand through 12 AI engines →
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.