Skip to content
AI VisibilityResearch
EN

We Asked the Same Category Two Ways. The Buying Question Cited Eight Sources, the Definition Cited None

Same category, same market, same day. Asked which platform to buy, ChatGPT cited eight domains. Asked what the thing is, it cited zero.

· Updated · 24 min read

On 8 August 2026 we put two questions about email marketing to ChatGPT and Gemini, minutes apart, in the same market and the same language. “What is the best email marketing platform for ecommerce?” came back with eight cited domains from ChatGPT and five from Gemini. “What is email marketing automation?” came back with zero from both. Same category. Same day. Same two surfaces. The only thing that changed was whether the question asked what to buy or what a word means.

That is the finding, and it has a blunt consequence: the “what is X” explainer, the single most recommended format in every generative engine optimization guide including several we have reviewed, is a format that two of the four surfaces we measure will often refuse to cite anyone for.

How this was measured: Six definitional questions and one buying question, run through the real interfaces of ChatGPT, Gemini, Google AI Overview and Google AI Mode on 8 August 2026 with country and language set explicitly on every call, using our own measurement tooling. We count cited domains only, never the results a surface searched and then ignored, because conflating those two is an error we have made in public before and corrected. Every run below is in our measurement register with its date and its surface count. The rows are in our measurement register, each with its surfaces and its run count.

Updated the same day with repeated, uncached runs. This article first published with one run per surface per question, and the fair criticism of that is that a single run is a draw and not a rate. We have since repeated the controlled pair on a path that bypasses caching. The finding held. The section on repetition below gives the numbers and, more usefully, explains why the obvious way of repeating a measurement in this category produces a fake denominator. The pair has since also been replicated in two unrelated categories, CRM and project management.


The short version

  1. The controlled pair is the whole study. Same category, same market, same day, same two surfaces: buying question, ChatGPT 8 domains and Gemini 5. Definitional question, both zero.
  2. Across six definitional questions in four languages, ChatGPT cited sources in exactly one. Across nine buying questions, it cited sources in all nine.
  3. Google’s two surfaces cite either way. AI Overview and AI Mode returned sources for every question of both kinds. The split is a chat assistant behaviour, not an AI search behaviour.
  4. The one exception is instructive and we could not reproduce it. The German question about LLMO is the only definitional question where ChatGPT cited anything, and the only one where it ran a search at all.
  5. It replicates in categories we have nothing to do with. Nine uncached runs across definitional questions in email marketing, CRM and project management returned zero cited sources. The buying question in each of the same three categories cited twelve to fifteen domains.
  6. It survived repetition, and the obvious way to repeat it is broken. Three uncached runs per chat surface kept the definitional zero. Repeating through a normal cached request returns byte-identical text even for a reworded question, so “ten runs” that way is one observation with a costume on.
  7. And ten days later the one citable definition stopped being citable. Re-run on 18 August 2026, the German LLMO question went from 3 cited domains and a recorded search to 0 and no search at all, while the buying question run minutes later cited 4 domains and recorded 4 fan-out queries.
  8. This corrects one of our own articles. Our German LLMO piece concluded the definition layer was the more stable place to compete. The measurement it rested on only sampled the two surfaces that always cite. Corrected below.

The controlled pair

Everything else here is breadth. This is the part that is actually controlled.

Question, US English, 8 August 2026ChatGPTGemini
what is the best email marketing platform for ecommerce?85
What is email marketing automation?00

Both questions are about email marketing. Both were asked in the United States in English. Both ran on the same day through the same two surfaces. One asks which product to choose and one asks what a term means.

On the buying question ChatGPT cited emailtooltester.com, a Contentful-hosted BigCommerce guide, and the homepages of Klaviyo, Omnisend, ActiveCampaign, Mailchimp, Shopify and Brevo. It also searched eleven further domains it did not cite, which is why we publish citations and not search lists. Gemini cited klaviyo.com, maropost.com, emailtooltester.com, insiderone.com and emailvendorselection.com. Two domains, emailtooltester.com and klaviyo.com, were cited by both.

On the definitional question, both surfaces produced long, competent, well-structured answers. ChatGPT’s ran to roughly five hundred words with a worked example, a diagram of a welcome sequence and a list of six named platforms. Gemini’s had headings, a three-part breakdown of triggers, workflows and content, and four worked examples. Neither cited a single source.

The answers were not worse. They were unattributable. There was no bibliography for any publisher to be in.

The same split across four languages

We ran the definitional form in four languages to see whether the pair above was a fluke of one topic.

Definitional questionMarketChatGPTGeminiAI OverviewAI Mode
What is generative engine optimization?US, English04810
What is generative engine optimization and how does it differ from SEO?US, English05not runnot run
What is email marketing automation?US, English00823
¿Qué es la optimización para motores generativos (GEO)?Spain, Spanish00910
Qu’est-ce que l’optimisation pour les moteurs génératifs (GEO) ?France, French00814
Was ist LLMO und wie unterscheidet es sich von GEO?Germany, German371211
ChatGPTGoogle AI ModeEmail automation, US0 to 23GEO, France, French0 to 14LLMO, Germany, German3 to 11GEO, US, English0 to 10GEO, Spain, Spanish0 to 10
The same definitional question, asked of a chat assistant and of a Google surface. ChatGPT cited nothing at all in four of these five, while AI Mode cited between ten and twenty-three domains for the identical question. Measured 8 August 2026, country and language set explicitly, uncached runs.
Distinct domains cited for definitional questions, ChatGPT against Google AI Mode
ChatGPTGoogle AI Mode
Email automation, US023
GEO, France, French014
LLMO, Germany, German311
GEO, US, English010
GEO, Spain, Spanish010

Read the ChatGPT column down. Five zeros and a three.

Now read the same column for the buying questions we have measured, all of them already published and registered: eleven for running shoes, sixteen for email platforms, four for the American GEO agency question, one for the French, three for the German, one for the second German wording, one for the Spanish, six for the American AI visibility tool question, and eight for the email platform question above. Nine buying questions, nine non-zero counts, ranging from one to sixteen.

Six definitional questions, one non-zero count.

One caveat we added the same day, because it cuts against the tidy version of this. A tenth buying question, “what is the best accounting software for freelancers?”, was measured uncached on ChatGPT after the above. Its first run cited nothing at all, and a second call of two runs cited twelve domains with every one appearing in exactly one of the two. So a buying question can return zero, and the same buying question can swing from zero to twelve between runs. The sentence above says nine buying questions gave nine non-zero counts, which is true of those nine and should not be read as “buying questions cite”.

The asymmetry survives that, and is worth restating precisely rather than conveniently. On the definitional side the zero held across nine uncached runs in three unrelated categories, with no run producing a single citation. On the buying side the count ranges from zero to sixteen, sometimes on the same question. Never citing and sometimes not citing are different claims, and only the first is ours.

Gemini splits the same way but less sharply: it cited on eight of nine buying questions and on three of six definitional ones.

The two Google surfaces do not split at all. AI Overview and AI Mode returned cited sources on every question of both kinds they answered, with one exception that is not a silence: on the American AI visibility tool question, AI Overview returned an upstream error instead of an answer, and that row is registered as three surfaces rather than counted as a zero. Everywhere else, both cited. If your definitional page is going to be cited anywhere, that is where.

We repeated it, because one run is a fair criticism

The most common objection to work like this, and the correct one, is that a single run per surface is a draw rather than a rate. Anyone can get a zero once. So we repeated the controlled pair. Getting that right turned out to be more interesting than the result.

The obvious method is void. We first re-ran “What is email marketing automation?” five times in a row. All five came back with zero citations, which looks like confirmation until you read the text: all five were identical, byte for byte, on both ChatGPT and Gemini. Then we changed the string. “What is email marketing automation” without the question mark returned the same text again. “Can you explain what email marketing automation is?”, which is a differently worded question, returned the same text a seventh time, down to the same hand-drawn ASCII diagram of a welcome sequence.

Three different inputs, one identical output. That is a cache, and it is not keyed on the exact string, because the three strings differ. Publishing “zero citations in ten runs” off the back of that would have been a fabricated denominator: it is one observation served seven times.

So we re-ran the pair on a path that bypasses the cache and re-plans retrieval on every request. That is the measurement below.

Question, US English, 8 August 2026SurfaceUncached runsDistinct domains cited
What is email marketing automation?ChatGPT30
What is email marketing automation?Gemini30
what is the best email marketing platform for ecommerce?ChatGPT215
what is the best email marketing platform for ecommerce?Gemini39

The definitional zero survives repetition. Three uncached runs on each chat surface, no citations in any of them.

The buying side produced the more surprising number. Across the two uncached ChatGPT runs, fifteen distinct domains were cited and every one of them appeared in exactly one of the two runs. The two runs shared no cited domain at all. Gemini was steadier but not steady: of nine domains, three appeared in all three runs (drip.com, maropost.com and emailtooltester.com) and six appeared in only one.

That asymmetry is the practical lesson, and it cuts against our own first draft as much as anyone’s. On a definitional question a single run is adequate, because zero has no variance and repeating it tells you nothing new. On a buying question a single run is close to worthless, because the citation set is substantially reshuffled between runs. Any vendor reporting your “AI visibility” for a commercial query off one daily run is reporting a coin flip, and any vendor repeating a query through a cache is drawing you a flat line by construction. Both are worth asking about before you buy.

Three categories, three pairs, nine uncached runs, no citations

The second fair objection, after the one about run counts, is that all of this is about generative engine optimization and email marketing, and we sell generative engine optimization. A finding that only shows up in your own industry is a finding about your own industry.

So we ran the same controlled pair in two categories we have nothing to do with, uncached, on ChatGPT.

CategoryDefinitional questionDomainsBuying questionDomains
Email marketingWhat is email marketing automation?0 in 3 runsbest email marketing platform for ecommerce?15 in 2 runs
CRMWhat is a CRM?0 in 3 runsbest CRM for a small business?12 in 1 run
Project managementWhat is project management software?0 in 3 runsbest project management software for a remote team?13 in 1 run

Nine uncached runs across three definitional questions in three unrelated categories, and not one cited source. Three buying questions in the same three categories, all cited, twelve to fifteen domains each.

The run counts are uneven and it is worth being explicit about which way that cuts. The definitional half has three uncached runs each because a zero is only worth anything if you repeated it. The buying half has one or two because ChatGPT is slow enough that the tool stopped early, and because that half only has to establish that citations happen at all, which one run returning thirteen domains does comfortably. If the asymmetry favoured our claim we would not publish it. It runs the other way: the half we repeated hardest is the half that returned nothing.

One detail from the buying runs is worth more than the counts. On the CRM question, three of the twelve cited domains are trade media you would expect, technologyadvice.com, forbes.com and techradar.com. The other nine include konabayev.com, softabase.com, themona.global, b2bbrief.com and americanbusinesscenter.org, which are not names anyone in the CRM market would recognise. On the project management question, six of the thirteen are the vendors themselves, monday.com, asana.com and clickup.com plus three of their support subdomains.

So the buying layer is reachable, and reachable by small sites. That is exactly what makes the definitional result worth acting on: it is not that citations are hard to win, it is that on these questions there are none to win.

The exception we could not explain away

The German LLMO question is the only definitional question where ChatGPT cited anything, and the obvious explanation is wrong.

Our first guess was the phrasing. That question asks for a comparison, “what is LLMO and how does it differ from GEO”, rather than a plain “what is X”, and a comparison plausibly needs reconciling two sources. So we tested it: we asked “What is generative engine optimization and how does it differ from SEO?” in the United States in English, same day, same two surfaces.

ChatGPT cited zero. The comparative form changed nothing. Gemini went from four cited domains to five, which is noise at this sample size.

So the comparative structure is not what did it. What is visible in the German run and absent from every English one is that ChatGPT actually searched. Its own fan-out query is recorded: “LLMO definition GEO generative engine optimization difference”. It searched nine domains and cited three of them, and its answer says out loud that the two terms are used “teilweise synonym”, partially synonymously. On the English questions there is no fan-out query and no search at all.

The reading that fits is that the assistant searches when it treats a term as unsettled and answers from its weights when it treats the term as settled. LLMO is a recent label with genuine disagreement about whether it differs from GEO. Email marketing automation is not in dispute. Generative engine optimization, in English, has a Wikipedia article and apparently counts as settled.

That is one observation and we are labelling it as one. It is a hypothesis with a visible mechanism, not a result. What we can say without hedging is the negative: the comparative phrasing did not reproduce the effect, so anyone planning to win ChatGPT citations by rewriting “what is X” as “what is X and how does it differ from Y” should measure that themselves before believing it.

Ten days later, the term we were watching stopped being citable

When this published we said the prompts were frozen and would be re-run, and we named the number to watch: whether the English GEO definition starts triggering a search again, or whether the German one stops.

On 18 August 2026 we re-ran them. The German one stopped.

Was ist LLMO und wie unterscheidet es sich von GEO?8 August18 August
ChatGPT, cited domains30
ChatGPT, search performedyesno
Gemini, cited domains74

On 8 August this was the only definitional question of the six where ChatGPT cited anything, and the only one where a fan-out query was recorded. Ten days later there is no fan-out query and no citation. The answer is still competent, still says the two terms are used “teilweise synonym”, and is now produced without looking anything up.

The control is what makes that readable. A definitional zero means nothing on a day when the tool is not retrieving at all, so we ran the buying question the same day, minutes apart: 4 cited domains and 4 fan-out queries recorded, against 45 search results. Retrieval was working. It just was not used on the definition.

That is the direction this article predicted, and it is worth being exact about what it is and is not. It is two points in time, not a series. It is one run per cell, and the re-read came back byte-identical, which is the cache this article already documents rather than a second observation. And an observed direction is not a demonstrated cause: nothing here shows that the term settling is what closed the window, only that the window closed in the ten days after we said to watch it.

The English GEO definition, the other half of the pair we named, timed out twice on the chat surface and is not measured here. It was zero on 8 August, so the interesting half was never it.

What it costs a plan, if it holds. The window to be cited for defining a category is not open indefinitely and may be measured in weeks. LLMO was a term young enough on 8 August that the assistant went and looked; ten days of the web writing about it was apparently enough for it to stop. Anyone planning to own a definition should assume the citable phase is the early one.

What this costs a content plan

Every guide to this discipline, ours included, tells you to write the self-contained definitional paragraph. That advice is not wrong, and the Princeton GEO-bench experiment over ten thousand queries is still the best evidence for the format: adding quotations raised a source’s share of the answer by 41%, statistics by 30% and cited sources by 27%, while keyword stuffing scored below doing nothing.

But format advice only pays where there is a bibliography. On five of the six definitional questions above, ChatGPT produced no bibliography for anyone. The most citable definitional paragraph ever written would have been cited zero times in those answers, because no page was.

So the honest version of the advice has a condition attached:

  • A definitional page is a bet on Google’s AI surfaces, not on the chat assistants. AI Overview and AI Mode cited on every definitional question we ran, between 8 and 23 domains each time. That is a real and winnable surface.
  • On a settled term, the chat assistants are answering from training, not retrieval. Nothing you publish this quarter reaches an answer that never performs a search. The lever there is being in the next training corpus and being described consistently everywhere, which is a twelve-month move, not a content-calendar move.
  • The buying question is where retrieval happens. Nine of nine. If you want a citation this month from a chat assistant, the page has to answer a question that has a purchase at the end of it.
  • Measure the question, not the topic. Two questions about email marketing, one day apart in the same market, produced thirteen citation slots and zero citation slots. Anyone reporting “AI visibility for email marketing” as a single number is averaging across that gap.

The correction to our own article

On 8 August 2026 we published a German piece on LLMO which concluded that definitional questions produce far more agreement between surfaces than buying questions, and that the definition layer is therefore the more stable place for a brand to compete. It reported seventeen cited domains with six of them appearing on both surfaces measured, more than a third, against near-zero agreement on the buying questions.

The arithmetic was right and the comparison was not sound, for a reason we flagged in that article as a limitation and then reasoned past anyway. That run measured two surfaces, AI Overview and AI Mode. The buying runs it was compared against measured four. Those two Google surfaces are precisely the two that cite on every question, so restricting to them selects for agreement. We compared a two-Google-surface measurement against four-surface measurements and read the difference as a property of the question.

With four more definitional questions measured, agreement between those same two Google surfaces runs like this: 29% for the English GEO definition, 19% for the Spanish, 11% for the email automation definition, 10% for the French, and the German 35%. The German figure is the top of the range, not a typical value.

The German article has been corrected to say this. Its practical advice about writing an extractable paragraph stands, because that rests on the Princeton work and on which pages the Google surfaces actually cite. Its claim that the definition layer is the more stable place to compete does not stand, and the replacement finding is the one in this article: on definitional questions the chat assistants often cite nobody, which is a harder problem than low agreement.

We also closed the gap that article declared. It said ChatGPT and Gemini were never asked its question because the four-surface call timed out. They have now been asked, on the same day, and the answer is three cited domains and seven.

What we did not measure

  • Repetition beyond the controlled pair. The pair was repeated on the uncached path, three runs per chat surface. The other five definitional questions and the four-surface rows are still one run each, and the register records the run count on every row for exactly this reason. Where a row says one run, it means one.
  • Perplexity and Claude. Our measurement reads ChatGPT, Gemini, Google AI Overview and Google AI Mode through their real interfaces. Those two are not in the set and we claim nothing about them.
  • Whether the uncited answers were correct. We counted who gets cited, not whether what they say is true. The uncited definitional answers looked accurate to us, which is part of what makes them hard to compete with.
  • Whether this holds as terms age. The mechanism we think we saw predicts that a term becomes uncitable as it settles. That is a claim about change over time and we have one day of data, so it is a question for a repeat run against the same frozen prompts, not a conclusion.
  • Logged-in or personalised behaviour. These are clean sessions. A user with history may get different retrieval decisions.

Where this goes next

The prompts above are frozen, and the first re-run against the same set has now happened, ten days later: it is the section above. The interesting number is not today’s split, it is whether the English GEO definition ever starts triggering a search again, or whether the German one stops. If the settled-versus-unsettled reading is right, a term should become less citable as it matures, which would mean the window to be cited for defining a category closes rather than opens.

If you want the same measurement on your own category, that is what our AI visibility product does: the same prompts, the same surfaces, country and language set explicitly, repeated so you get a rate instead of a draw. The full cross-surface study this one extends is four surfaces, one question, and the earlier and smaller observation that first raised this is answers with no sources.

Common Questions About Uncited Definitions

Does ChatGPT ever cite sources for a “what is” question?

Yes, but rarely in our measurement. Across six definitional questions in four languages on 8 August 2026 it cited sources in one, the German question about LLMO, where three domains were cited. In the other five it produced complete answers with no sources at all. Across nine buying questions measured on 7 and 8 August 2026 it cited sources in all nine. A tenth, measured later the same day, cited nothing on its first uncached run and twelve domains across two more, so buying questions are not guaranteed to cite either. What separates them is that the definitional zero held in every one of nine uncached runs, while the buying count ranges from zero to sixteen.

Why would an AI answer a definition without citing anything?

Because it does not search. In the runs where ChatGPT cited nothing there is no fan-out query recorded, meaning no retrieval was performed and the answer came from the model’s training. In the one German run where it did cite, a search query is recorded. A model that considers a term settled has no reason to look it up.

Does this mean definitional content is worthless for AI visibility?

No, it means the surface matters. Google AI Overview and Google AI Mode cited sources on every definitional question we ran, between 8 and 23 domains each. A definitional page is a bet on those surfaces. What it is not is a reliable way to be cited by a chat assistant on a term the model treats as settled.

No. Zero-click describes a user getting an answer without visiting a site, but the sources are still shown and attributed. This is stronger: there is no attribution at all, so there is no link for anyone to not click and no publisher named. A brand cannot appear in an answer that names nobody.

How many runs is this based on?

The controlled pair was repeated three times per chat surface on a cache-bypassing path, and the definitional zero held in every run. The other five definitional questions and the four-surface rows are one run each, and the register records the count on every row. A warning for anyone repeating this: re-asking the same question through a normal cached request returns identical text, so ten such “runs” are one observation, not ten. We hit that and say so above rather than publishing the inflated denominator it would have produced.

Is this only true in your own industry?

No. The controlled pair was repeated in CRM and project management, categories we have no involvement in. “What is a CRM?” and “What is project management software?” each returned zero cited domains across three uncached ChatGPT runs, while “what is the best CRM for a small business?” and “what is the best project management software for a remote team?” cited twelve and thirteen domains. Nine uncached definitional runs across three unrelated categories, no citations in any of them.

Does repeating the question change the answer?

It depends entirely on the question type, which is the useful part. On the definitional question, three uncached runs per surface produced zero citations every time. On the buying question, two uncached ChatGPT runs cited fifteen domains between them and shared none: every domain appeared in exactly one of the two runs. Gemini cited nine across three runs, of which three appeared every time. So a single run is adequate for a definitional question and close to meaningless for a commercial one.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)