Skip to content
AI VisibilityGEO
EN

Twenty Buying Questions in Four Markets, and One Category Where the Overview Barely Renders

Twenty buying questions in four markets, one reading each, no retries. Seventeen answered every run. All three that did not are the same category.

· 19 min read

Every AI visibility product reports a zero the same way. It cannot tell you whether the engine answered and left you out, or whether the engine did not answer at all. We asked buying questions in the United States, Spain, Germany and France, one reading per cell, three runs each, no retries whatever came back, and the study has grown to 84 readings across 10 categories, 5 question shapes and 2 surfaces. 37 of the 40 cells returned an AI Overview on every run. All 3 that did not are the same category, and in that category the surface returned 3 runs of 3 in Germany, 2 in the United States, 1 in France and 0 in Spain.

Whether the overview renders is far more a property of the question than of the market. But for the question where it is unreliable, it is unreliable differently in every market we asked, which is the part a dashboard cannot show you.

Disclosure: EchoWi sells AI visibility measurement, and a study arguing that some visibility numbers describe a surface that did not run is a study with an interest. Everything needed to repeat it is here: every exact string, the surfaces, the four markets in their own languages, three runs per cell, the dates, and all 84 readings in our measurement register. Each arm is twenty calls, so anyone can run any of them against us.


The short version

  1. 40 cells: 10 buying questions in 4 markets, each in its own language. One reading each, three runs requested, on Google’s AI Overview.
  2. 37 of 40 returned an overview on all three runs. So for most buying questions in most markets, the surface is simply there.
  3. All 3 shortfalls are one category, password managers. Every other cell in every market answered fully.
  4. That one category walks the whole range. Germany 3 of 3, the United States 2 of 3, France 1 of 3, Spain 0 of 3, all on the same night.
  5. No retries, deliberately. A short reading is the result here, not a failure to be corrected, and re-reading only the short ones is how a sample gets selected by its own outcome.
  6. The prediction we wrote down first was wrong, and the section on it says how.

Why anyone should care what the denominator is

A visibility report tells you that you were not cited. There are two completely different worlds behind that sentence.

In the first, the engine produced an answer, chose eight sources, and none of them was you. That is a content and authority problem, it is the one every tool in this category is built to describe, and the work that follows from it is the work we sell.

In the second, the engine produced nothing. There was no answer, no citation list, and no eight sources. Nobody was cited, including your competitors, including the incumbents. Your zero is identical to everybody’s zero and it says nothing at all about your pages.

The two look the same in every dashboard we have seen, ours included. The number is zero either way.

We have arrived at this distinction twice from different directions. A sweep of local buying questions found a whole question shape where Google’s AI Overview never ran at all, while AI Mode answered every time on the same call. And in a study of keyword phrasing against question phrasing, one French keyword phrase returned nothing across six runs while the question form of the same intent rendered and cited five stable sources. In both cases the interesting fact was not who was cited. It was that there was nothing to be cited in.

What neither of those told us is how often that happens for ordinary buying questions, which is what most measurement is actually pointed at.

The design, and the one rule that makes it worth running

Five categories, chosen because two earlier arms of our keyword study already used them, so the strings are ones we have measured before and can be checked against the register. Each asked in four markets, in that market’s own language, using the keyword form so that phrasing is held constant. Three runs per cell on the AI Overview. Twenty cells in total.

One reading per cell, and no retries, whatever came back. That rule is the study.

Counting answered runs across every AI Overview reading already in our register produces a tidy-looking market gradient. It is worthless, for three separate reasons, and each one is enough on its own. The arms cover different question sets, nineteen categories in one market and five in the others. They were read on different days. And short readings were re-read, so the arm that produced more short readings collected more extra rows, which means the sample was selected by exactly the outcome being measured.

That last one is the trap this register has already documented in its own control numbers, and it is worth naming plainly: if you retry the readings that came back thin, your retry policy becomes your result.

So no retries. A cell that answered once out of three is recorded as one out of three and stays there.

What came back

CategoryUnited StatesSpainGermanyFrance
CRM3333
Password managers2031
Web hosting3333
Online shops3333
Business banking3333
Answered of 1514121513

Sixteen of these twenty cells sit at three of three. The seventeenth full cell is German password managers, which is the point of the row below it.

Every shortfall in the study is one category. Four markets, five categories, and the only cell type that ever failed to render on every run is password managers.

That category then does something none of the others do. It returns a full three of three in Germany, two of three in the United States, one of three in France, and nothing at all in Spain, on the same night, with the same tool, asking the equivalent phrase in each language. The other four categories are flat across all four markets.

Five categories was not enough categories, so we added five more

“All three shortfalls are one category” rests on however many categories you asked. With five, it can be true because there is one badly behaved category in the world, or because five is not enough to find the second one. The study could not tell those apart and neither could a reader.

So five more went out: project management, email marketing, accounting, VPNs and website builders, in the same four markets, same rule, no retries. The threshold was written first: nothing new short and the claim gets published over ten categories; anything new short and the headline is rewritten in four languages to name the count of categories with a shortfall.

All 20 of the new cells answered every run. So the study is 37 of 40 across 10 categories, and every shortfall in it is still one category out of ten.

We are not going to explain that, because we did not measure a mechanism. What we can say is what it does to a number a buyer is being sold: in Spain, a report on that category built from a single reading would have shown zero citations for everyone in the market, and the correct reading of that zero is that there was nothing to be cited in.

We read the one interesting cell again, before writing about it

A single reading that produces a clean gradient is the shape this register has been burned by. The rule we wrote for ourselves is that when one cell gives a dramatic result, the next call is the replication and not the write-up. So the password-manager question was asked a second time in all four markets, same strings, same three runs, same no-retry rule.

Password managersFirst readingSecond reading
Germany33
United States23
France11
Spain00

The ordering held and nobody crossed anybody. Germany is still at the top, Spain is still at the bottom, and the one cell that moved moved upward: the United States went from two of three to three of three, which ties it with Germany rather than reordering anything.

Two things are worth more than the ordering.

Spain is now 0 of 6 answered runs across two readings on two dates. Six requests for an AI Overview on an ordinary commercial buying question, and six times nothing rendered, while the equivalent German phrase rendered on all six of its. A visibility report for that category in that market, built from any number of runs, would show every vendor at zero.

France returned exactly one of three, twice. That is a different thing from a cell that is merely noisy. Two independent readings landing on the same fraction looks like a partial render rate rather than a bad night, and it is the first cell in this study that can be described that way.

Both readings are printed and neither replaces the other. A mean of two would hide the only two numbers a reader can check, and with counts this small the mean is the least informative thing available.

The headline count of 17 of 20 is deliberately still the first reading of each cell. Folding a replication into it would mean the number changes whenever we re-read something, which is how a figure stops being checkable.

Four other things a buyer types

Every cell so far is the same shape, best X. That is one of several things someone types before buying. They also ask what it costs, how the options compare, how to choose, and whether there is a free one, and a report that says a category has no visibility is pooling all of them.

So the five original categories were asked four more ways, in the United States, same rule and no retries, holding market and category fixed so shape is the only thing moving. The best row already existed, which makes this a comparison within category as well as within market.

Answered of 3CRMPasswordHostingEcommerceBanking
best32333
price32303
comparison31333
how to choose13333
free33333

16 of the 20 new cells answered every run. The 4 that did not split into two kinds, and keeping them apart is the whole point of this study: 2 came back short because the surface did not render, and 2 because the tool stopped early. The zero in the ecommerce price cell is the second kind, not the first.

Both of the surface shortfalls are password managers. Counting its best cell too, that category is short in 3 of its 5 shapes while no other category is short anywhere the measurement completed.

So shape does not appear to be what decides whether an overview renders, and the category reading survives being asked four more ways. That is worth more than the original arm, because “one category is unreliable” measured five ways is a different claim from the same sentence measured one way.

One market, and this section says nothing about the other three. The shapes were asked in English in the United States only.

The same twenty questions on the other surface

The arm above cannot say why a cell came back short, and it does not try. There is one thing it can say cheaply that changes the advice: whether the other surface answers the same question.

So the same 20 cells went to AI Mode. Same strings, same markets and languages, same three runs, same no-retry rule.

Answered of 15AI OverviewAI Mode
United States1415
Spain1215
Germany1515
France1314

Every cell that came back short on the AI Overview returned a full three of three on AI Mode. All 3 of them. Paired within the cell, AI Mode was higher in 3 cells, equal in 16, and lower in 1.

The sharpest case is the one this study has been circling. The Spanish password-manager phrase has now gone 0 of 6 requested AI Overview runs across the two readings of this study. On AI Mode, the same string in the same market on the same day answered 3 of 3 and cited sixteen domains.

That settles what the Spanish zero is. It is not a question nobody can answer, and it is not a category with no sources. It is one surface deciding not to render, while the other one answers the same question and hands out sixteen citations. A visibility report showing zero for that market and category is describing Google’s layout, not anybody’s content.

The threshold this arm was frozen with asked for 20 of 20 and it returned 19. The exception is French web hosting at two of three, and its result carries the tool’s own early-stop flag, so what ran short was the measurement rather than the surface declining. We are treating that as the first branch with the exception named rather than as a third pattern, and this sentence exists so a reader can disagree with the call.

One thing this cannot do, and it is the reason there is no combined percentage anywhere above: the two arms were read on different days. Every comparison here is within a cell, where the question and the market cancel out, and none of it is a rate across surfaces.

The prediction we wrote down first, and why we were wrong

The design was frozen before the first call, with the prediction and the reporting threshold written into it. Both are worth quoting because the prediction failed.

We predicted that the ordering of the earlier, contaminated gradient would roughly reproduce: the United States highest and France lowest. The threshold said that if it did not reproduce, we would publish that it did not, and correct the earlier observation rather than soften it.

Germany answered every requested run and Spain answered fewest. France ties within one run of the United States. Neither end of the prediction survived.

So the earlier gradient was the question sets, the different days and the retry selection, not the markets. An instrument note in our keyword study said France delivered fewer full readings than Germany, and with the selection removed France is within one run of the American arm and Germany is perfect. That note is corrected in all four of its languages rather than quietly dropped, and the correction is here so a reader who saw the first version can find the second.

There is a smaller thing worth admitting alongside it. The same French password-manager phrase returned nothing across six runs in two readings the day before, and one of three here. Three readings of one string across two days: nothing, nothing, then something. A single reading of that cell could honestly have reported either “the surface never runs here” or “the surface runs a third of the time”, and those are different claims.

What this changes if you are buying or building measurement

Ask any vendor what their zero means. Not whether they measure AI Overviews, which everyone says yes to. Ask whether a zero in their report distinguishes “the engine answered and did not cite you” from “the engine did not answer”. If they cannot answer that, the number is two numbers wearing one label.

Ask how many runs are behind a cell, and what happens to the short ones. A single run is a coin toss dressed as a measurement, and a tool that silently retries until it gets an answer is reporting its retry policy.

Do not generalise a render failure across markets. The one category that failed here failed by a different amount in every market, and the four categories that did not fail did not fail anywhere. Whatever governs it is not a national setting you can look up.

And do not generalise it across surfaces either. This is Google’s AI Overview only. Our own work has repeatedly found that a result measured on one surface does not travel to another, and the local sweep is the sharpest example: the surface that refused to answer at all was answered every time by Google’s own AI Mode, on the same call.

The practical version, for anyone running their own checks: record how many runs answered, next to the citation count, every time. It costs nothing, it is the difference between the two zeros, and almost nobody does it.


Limits

40 cells. Ten categories in four markets is enough to show that both outcomes exist and nowhere near enough to estimate how often either occurs. Treat every count here as a count.

One reading per cell, except one. That is the design and it is also the main limit: a single reading cannot separate a category that renders a third of the time from one that happened to be measured on a bad night. The one cell that produced a result worth reporting was read a second time in all four markets, and those 4 readings are reported beside the first rather than folded into it. Every other cell in the study is still a single reading.

2 surfaces, and only one of them everywhere. Google’s AI Overview across every arm, and AI Mode on the twenty cells of the paired section. Nothing here describes ChatGPT, Gemini or Perplexity.

5 question shapes, and only one of them in every market. The keyword phrase is held constant across all four markets so that markets can be compared; the other four shapes were asked in the United States only. We already know that the question form of a French password-manager query rendered when the keyword form did not, so form matters and this study now measures some of it rather than only holding it fixed.

2 calendar dates. The first session crossed local midnight, so its rows carry two dates for under forty minutes of elapsed time. Everything added afterwards was read on the second of those dates. No arm here spans a day.

Nothing here is causal. No page was changed and nothing was measured before and after. This is what the surface returned.


Common Questions About Whether the AI Overview Renders

No. In this study 37 of 40 cells returned one on every run, so it usually does for buying questions, but one category returned nothing at all in one market and one run of three in another. A separate sweep of local service questions found a whole question shape where it never appeared. So “usually” is the right word and “always” is not.

If the AI Overview does not render, does that mean I am not visible?

It means the opposite of what a dashboard implies. If the overview did not render, nobody was cited in it: not you, not your competitors, not the incumbents. A zero from a run where nothing rendered says nothing whatsoever about your pages, while a zero from a run where the engine answered and chose eight other sources says a great deal. They are the same number and they call for opposite responses.

Does the AI Overview render less often outside the United States?

Not in this study, and we predicted that it would. Germany returned an overview on all 30 requested runs, the United States on 29, France on 28 and Spain on 27. The German arm was the most complete and the American arm was not the top. An earlier count of ours suggested otherwise and was measuring its own retry policy.

Why does one category behave differently from all the others?

We do not know, and we did not measure a mechanism, so we are not going to offer one. What we can say is that it is not a market setting: the same category returned three of three, two of three, one of three and none of three in four different markets on the same night, while the other four categories returned three of three everywhere.

How many runs should a visibility measurement use?

More than one, and the number should be visible in the report. A single run cannot tell a stable citation from a coincidence, and it cannot tell an absent citation from an absent answer. Three is the minimum we use for anything published, and even three is thin: a strict “cited in every run” bar drops a source entirely if one run goes astray.


Where this leaves you

The useful takeaway is not that AI Overviews are unreliable. For four of the five categories here they rendered on every run in every market, which is a stable surface by any reasonable standard.

The takeaway is that the reliability is not uniform, it does not follow the market, and the failure mode is invisible in the only number most tools report. One category in this small study produced a zero in Spain that means nothing about anybody’s content, and it would have appeared in a report identical to a zero that means everything.

If you take one habit from this, take the cheap one: count the answers, not only the citations.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)