What Is GEO? Generative Engine Optimization vs SEO
GEO is how brands get cited by ChatGPT, Perplexity and AI Overviews. What the research proves, what it does not, and how to optimize a page.
What is GEO (Generative Engine Optimization)?
GEO is the practice of making a brand findable, usable and citable inside AI-generated answers. Where SEO competes for a position in a list of links, GEO competes to be the source a model draws on when it writes a recommendation.
The term comes from a 2023 paper by researchers at Princeton, Georgia Tech, the Allen Institute and IIT Delhi, who built a benchmark of 10,000 queries to test whether editing a source document changed how much a model used it. It did. That paper is still the foundation of the field, and its limits are still the field’s blind spot.
In one sentence: GEO is optimizing so that AI systems can retrieve your page, understand it, and attribute a claim to you by name.
The short version
- 51% of B2B software buyers now begin a purchase inside an AI chatbot rather than a search engine, up from 29% a year earlier (G2, 2026).
- The click is disappearing faster than the visit. When Google shows an AI summary, 8% of visits end in a click, against 15% without one (Pew Research Center, 2025).
- What the evidence supports is unglamorous: be retrievable, be relevant to the sub-questions a model actually generates, and carry quotable evidence with a name on it.
- Several popular tactics have been tested and failed. Adding schema to already-cited pages did not increase citations. Keyword stuffing scored below doing nothing.
- Language matters more than most teams assume. For French prompts, 72% to 87% of cited sources were themselves in French, depending on the engine.
One caveat runs through this whole article. Almost everything published about GEO measures what a model does with a document it has already been handed. Very little measures whether your page gets retrieved from the open web in the first place. Those are different problems, and confusing them is how this field produces confident advice that does not survive a control group.
How a source ends up in an AI answer
A citation is the last step in a chain, and most GEO advice only addresses the last two links.
- Crawl. The engine’s crawler has to be allowed in. To appear in ChatGPT Search that means OAI-SearchBot, not only Googlebot (OpenAI documentation). Allowed in
robots.txtis not the same as allowed at the edge, and the difference is invisible to the test everyone runs: a CDN can be challenging one crawler while every other one passes. - Index. The page has to be stored and rendered. Content locked behind heavy JavaScript, or sitting inside an image, is content that does not exist.
- Query fan-out. This is the step teams miss. The model does not search for your prompt. It breaks it into several sub-queries and searches for those. Ahrefs found that cited URLs had visibly higher semantic similarity between their titles and those derived sub-queries, not the original prompt.
- Rerank and select. A handful of candidates make it into the model’s context window.
- Generate and attribute. The model writes an answer and decides which of those candidates to lean on and name.
Steps 1 to 4 are search and retrieval engineering. Only step 5 is where GEO copywriting applies. A page written perfectly for step 5 that fails step 1 gets nothing.
This also settles a common worry. You do not need to rank in the top 10 for the literal prompt to be cited. Engines run sub-queries and pull URLs that never ranked for the original phrasing.
GEO vs SEO: what actually differs
| SEO | GEO | |
|---|---|---|
| Competes for | A position in a list of links | A sentence inside a generated answer |
| Unit of success | Ranking, then a click | Being retrieved, then being cited by name |
| Who sees it | Anyone who scrolls | Only whoever the model names |
| Input signal | Keywords, chosen for volume | Conversational prompts, and the sub-questions a model derives from them |
| Authority signal | Backlinks and domain metrics | Consistent entity data, extractable structure, being cited elsewhere |
| Typical metric | Position, impressions, CTR | Mention rate, share of voice, citation frequency, sentiment |
| Cost of losing | Page two | Absence, with no page two to be on |
| Relationship | The foundation | Built on top of it, not instead of it |
They are not rivals. Retrieval runs on a search index. A page that cannot be crawled, or that ranks for no sub-query at all, is not a GEO candidate in the first place. GEO is what you do after SEO works, not instead of it.
What the research actually shows
This is the part most articles skip. Below is every substantial study we could verify, with what it measured and what it does not prove.
| Study | Scale | What it measured | Headline result | The limitation that matters |
|---|---|---|---|---|
| Aggarwal et al., GEO-bench, 2023 | 10,000 queries | How much of an answer came from an edited source already in context | Adding quotations +41%, statistics +30%, cited sources +27% relative. Keyword stuffing scored below baseline | The source was already retrieved. Says nothing about being found |
| Vishwakarma et al., “What Gets Cited”, SIGIR 2026 | 252,000 controlled comparisons across 6 models | Which of two documents got the first citation | Topical relevance dominated. Explicit price and recency helped slightly. Pure formatting changes did almost nothing | Both documents were injected into context. Post-retrieval again |
| Ahrefs, why ChatGPT cites pages | 1.4M prompts, 46.8M candidate URLs | Which retrieved URLs got cited | Similarity between page title and the model’s derived sub-queries separated cited from uncited | Observational, and similarity was approximated with open embeddings |
| Ahrefs, schema and AI citations | 1,885 pages plus ~4,000 controls | Citations 30 days before and after adding schema | AI Overviews −4.6%. AI Mode +2.4%, not significant. ChatGPT +2.2%, not significant | Not randomised, and only on pages already heavily cited |
| Semrush, content and AI search | 304,805 cited URLs vs 921,614 ranking URLs | Text qualities of cited pages | Clarity and summarisation +32.8%, E-E-A-T signals +30.6%, Q&A format +25.4% | Correlational. Cited pages differ in many ways at once |
| Temso, local language sources | 7M citations, 6 languages, 12 countries | Language of the sources cited, per prompt language | French prompts: AI Overview 82.1%, Copilot 87.2%, ChatGPT 72.3%, Grok 58.9% French-language sources | Vendor study, not peer reviewed. Content availability explains part of it |
A July 2026 review of 45 GEO studies reached a blunt conclusion: no technique has yet demonstrated a causal, stable, longitudinal, cross-platform effect on organic retrieval or on business outcomes (Martinez, 2026).
That is not a reason to do nothing. It is a reason to treat these as predictors, and to measure your own results against a control, rather than promising an uplift nobody has proven.
The signals that look real
Ordered by strength of evidence, not by how well they present in a pitch deck.
-
Relevance to the derived sub-questions. The strongest and most boring finding, and it shows up in every study. Write sections that answer the questions a model would generate from a prompt in your category, in the words a buyer would use.
-
Quotable, attributable evidence. The three winning edits in the Princeton experiment were quotations, statistics and cited sources. All three share one property: they give the model something it can lift and attribute to you by name. A paragraph of adjectives gives it nothing to hold.
-
Self-contained passages. A model lifts passages, not pages. A section that only makes sense after reading the previous three is a section that cannot be quoted.
-
Authority and outside corroboration. Links, editorial coverage, presence in the sources of your sector. Weaker and slower than vendors claim, but it decides whether you are a candidate at all.
-
Accessibility. Crawlable, renderable, allowed. This one is binary: fail it and nothing else counts.
-
Language. Covered below, and with stronger evidence than most teams expect.
What is not proven
Being straight about this is itself a trust signal, and it is the part of GEO advice most likely to waste a budget.
- “Adding schema increases AI citations.” Tested with a control group. It did not, on already-cited pages. Use structured data because it describes your page correctly and feeds search features, not as a GEO lever.
- “The Princeton paper proved a 40% lift in citations.” It proved a 41% relative lift in how much of an answer came from a source already sitting in the model’s context. Very different claims.
- “LLMs prefer 40 to 60 word blocks.” No primary evidence for that range exists.
- “The ideal length is 1,500, or 2,000, or 3,000 words.” Not demonstrated. Length should follow the question.
- “llms.txt improves visibility.” Google has stated it neither helps nor hurts rankings or visibility.
- “Every H2 should be a question.” It may help readers and topic coverage. No isolated causal effect has been shown.
- “Refreshing the date improves GEO.” Not without a substantive update, and Google advises against it explicitly.
- “More backlinks means more ChatGPT citations.” Links matter to the search layer underneath. The direct association in brand studies is weak.
- “One prompt check proves the optimization worked.” Month to month source drift can exceed 50%. Without repetition and a control, you are measuring noise.
How to optimize a page for GEO
- Map the derived sub-questions, not the head keyword. If someone asks an assistant for the best tool in your category, what four or five things does it need to look up to answer well?
- Answer each one in its own section, with the thesis in the first sentence. Assume the reader arrives mid-page, because the model does.
- Put a verifiable fact in every section: a number, a date, a named source, a direct quotation. This is the single edit with controlled experimental support behind it.
- Attribute everything. “Studies show” is unusable to a model. “Pew Research Center, 2025, measuring 68,879 searches” is quotable.
- Keep passages self-contained. If a paragraph needs three previous paragraphs to make sense, it will not survive extraction.
- Link to it internally from related pages, so retrieval has a path in.
- Check the plumbing: crawlable, renderable, in the sitemap, correct canonical, correct hreflang, and the AI crawlers you want allowed in
robots.txt. - Then measure against a control. A fixed prompt set, the same engine, model, country and language, several runs, and a comparable page you did not touch.
Differences by engine and by language
Engines do not cite alike. ChatGPT leans on general web search far more than on any single social feed: in the Ahrefs data the general search channel produced an 88.5% citation rate, against 1.9% for the dedicated Reddit feed. Perplexity is built on retrieval-augmented generation and cites more visibly. Google AI Overviews and AI Mode sit on top of the Google index, so classic SEO carries further there.
Language is the finding most teams underestimate. In Temso’s analysis of more than seven million citations, French prompts were answered mostly with French-language sources: 82.1% in AI Overview, 87.2% in Copilot, 72.3% in ChatGPT. Academic work on multilingual retrieval finds the same pull toward the language of the query (mRAG study).
That is observational, not an A/B test of a page against its own translation. But if a market matters to you, the practical reading is clear: publish a real version in that language, with its own URL, correct hreflang, its own examples and sources, and a canonical pointing at itself. A machine translation with an English canonical is not that.
How to measure whether GEO is working
You cannot manage what you measure once. A defensible measurement has five parts:
| Element | Why it matters |
|---|---|
| A frozen prompt set | Change the prompts and you change the result |
| Engine, model, country and language recorded | The same prompt answers differently across all four |
| Several runs per prompt | Answers vary run to run. One check is an anecdote |
| A control page you did not touch | Without it you cannot separate your work from the platform’s own drift |
| Mention, position, sentiment and sources | Appearing is not the same as appearing well, or being attributed correctly |
This is what EchoWi automates: a stable prompt set run on a schedule across 12 engines, with before and after measurement for every change, so an improvement can be attributed to something you actually did.
If you would rather compare before buying anything, our verified catalogue carries 64 tools with every price read from the vendor’s own page on a stated date. We are in it, on the same axes as everyone else, including the row where our entry plan covers one engine and a cheaper rival covers ten.
Common Questions About GEO
Is GEO just SEO with a new name?
No, but it depends on SEO completely. SEO decides whether your page can be retrieved. GEO decides whether, once retrieved, it gets used and named. Skipping the first and doing the second is optimizing a page nobody will ever fetch.
How long does GEO take to show results?
There is no established figure, and anyone quoting a precise one is guessing. It depends on how often the engine refreshes its index and its training data, and the only way to know for your site is a fixed prompt set measured against a control over several weeks.
Can small businesses compete in GEO?
In an established English-language category, competing on authority alone is hard. In an emerging category, or in a language where little quality content exists, the entry cost is far lower. That is why the language question above is strategic rather than cosmetic.
Does GEO work for B2B companies?
The strongest available evidence is B2B. G2’s 2026 survey of 1,076 software buyers found 51% now start inside an AI chatbot, 85% think more highly of a vendor the AI names, and 33% bought from a vendor they had never heard of before.
Which AI engine should I optimize for first?
Whichever your buyers use, which is worth measuring rather than assuming. Optimizing for retrieval quality tends to help across engines, because the underlying problem, being findable and quotable, is shared.
How is GEO different from AEO?
AEO, Answer Engine Optimization, generally refers to appearing in direct answer features such as featured snippets and assistant replies. GEO is broader: it covers how a generated answer is composed and which sources it leans on. The terms are still settling and are often used interchangeably.
Do I need to add schema markup for GEO?
Add it because it describes your page accurately and feeds search features. Do not add it expecting more AI citations. The one study that tested this with a control group found no significant increase, and a small decrease in AI Overviews.
Can I get alerts when an AI describes my brand wrongly?
Monitoring tools can detect it; none can remove it. What a platform can do is run a fixed set of prompts on a schedule, record the answer text and the cited sources each time, and flag when the description changes or a false claim appears. What follows is not a takedown, because there is no index entry to correct. It is publishing a correction on a page the engine already cites and waiting for the next crawl or refresh, which is why the cited-source list matters more than the sentiment score: it tells you which page to fix.
Can a platform show which prompts pick a competitor instead of me?
Yes, and it is the more useful half of the measurement. It requires storing two outcomes separately for every run: whether your brand was named, and whether your own domain was cited. A prompt where a competitor is named and their site is cited is a content problem on a page you can identify. A prompt where a competitor is named and no source is cited at all is a different problem entirely, because there is no bibliography to enter. Any tool reporting one blended visibility percentage has thrown that distinction away.
Do GEO platforms collect user data ethically, without PII?
The honest answer is that this class of tool does not usually touch user data at all. It sends your prompts to public assistant interfaces and records what comes back, so the data collected is answers about your brand rather than anything about your customers. The questions worth asking a vendor are narrower and rarely volunteered: whether runs bypass caching, since a cached repeat is the same observation served twice rather than a trend, and whether the country and language are set explicitly on every call, since most tools in this category default to the United States and English without saying so.
What are the best GEO practices for ChatGPT and Gemini specifically?
Distinguish the question type before the engine, because it decides whether a citation exists to win. Measured on 8 August 2026 across three unrelated categories, definitional questions such as “what is a CRM” or “what is email marketing automation” returned zero cited sources from ChatGPT in nine uncached runs, while the buying question in each of the same categories cited twelve to fifteen domains. Google’s AI Overview and AI Mode cited sources on both kinds. So a “what is X” explainer is a bet on the Google surfaces, and a page that answers a purchase question is what earns a chat assistant citation. Full method in the study.
Does my brand need to be in the training data?
Not necessarily. Retrieval-based engines can cite a page published this week. Training data helps a model know your brand without looking it up, which is a different and slower advantage.
Where this leaves you
The category is young enough that the honest position is also the differentiating one. Most GEO advice available today is folklore with a number attached. The research that does exist is real, and it points somewhere unfashionable: be retrievable, answer the question that was actually asked, and put a quotable, attributable fact in front of the model.
The brands treating this as measurement rather than ritual will find out what works for them while their competitors are still arguing about word counts.
Next: how brands get positioned inside LLMs, and how AI is changing the way people buy.
See what 12 AI engines say about your brand today →
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.