Skip to content
AI VisibilityResearch
EN

We Asked Google the Same Question Twice. In the US Nothing Changed and in France Half of It Did.

One buying question, asked twice minutes apart, in four markets. The sources that hold repeated 100% of the time in the United States and 50% in France.

· 14 min read

Ask Google one buying question, write down the sources its AI Overview cites on every run, then ask the identical string again a few minutes later. In the United States you get the same list back, name for name. In France you get half of it.

That gap is not about the question, because it is the same question. It is not about the day, because both readings are minutes apart. It is a property of the market you are measuring in, and it decides whether a single reading of your visibility means anything at all.

How this was measured: One bare buying question per market, localised, run on Google’s AI Overview through our own measurement tools. Every reading takes the domains cited on every run, then the identical string is asked again in the same session and the two sets are compared. All readings are at a matched two-run bar, so no comparison here puts a stricter set against a looser one. Four markets, three categories, 18 August 2026, with a second arm repeating six cells at three-, four- and five-run bars. The full rows are in our measurement register, and the exact prompts are below, so anyone can repeat this against us.

Disclosure: EchoWi sells AI visibility measurement, so a finding about how unreliable a single reading can be is a finding about the thing we sell. The prompts, the surface, the markets, the bar and the date are all here, which is what makes it checkable rather than something you have to take from us.


The short version

  1. The same question, asked twice, returns the same durable sources 100% of the time in the United States and 50% in France.
  2. Every market has a stable core. Between 3 and 6 domains come back every time you ask, in all four markets.
  3. What differs is the rim. The number of sources that appear in one reading and not the other runs from 0 in the United States to 4 in France.
  4. The market is most of it, and a middle market shows it is not all of it. Across 3 categories in the two extreme markets the French readings never rise above the American ones. But Germany returns 67% on CRM and 100% on accounting, both at matched bars, so inside one market the question moves the figure by a third.
  5. The practical consequence is a rule for reading any AI visibility report: if you are in the core you are in it every time, and if you are on the rim a single reading is a coin flip.
  6. A perfect reading expires. The American CRM question returned 4 durable domains earlier the same day and 12 later, at the same five-run bar, with every individual reading flawless.

What we measured, exactly

The unit is one bare buying question of the form “best X for a small business”, written the way a buyer in that market would write it, in that market’s language.

For each reading, the surface is asked twice and we keep only the domains cited both times. That set is the reading. Then the identical string is asked again in the same session, producing a second set, and the two are compared with a Jaccard: the domains in both, divided by the domains in either.

Two details do most of the work here and neither is decoration.

Every reading sits at the same two-run bar. A set of domains that survived five runs is a smaller and tougher set than one that survived two, so comparing across bars measures the bar as well as the world. An earlier version of this measurement had exactly that problem, with one market compared at five runs against two and another at four against three, and it is why those numbers are not in this article.

The prompts are frozen and published. They are:

MarketCategoryPrompt
United StatesCRMwhat is the best CRM for a small business?
GermanyCRMbeste CRM Software für kleine Unternehmen
SpainCRMmejor CRM para una pequeña empresa
FranceCRMmeilleur CRM pour une petite entreprise

The result

MarketComes backCoreRim
United States100%40
Germany86%61
Spain83%51
France50%44

That is 57 points between the widest and the narrowest. Before running the missing readings we wrote down a threshold: if all four markets landed within fifteen points of each other, there was no market difference worth reporting and the observation was to be dropped. They did not.

The American cell is the one worth pausing on. The same four domains came back on three separate readings, one of which was taken at a five-run bar rather than a two-run one. That is not a stable-ish result, it is the identical list every single time we asked.

France was the dramatic cell, so we replicated it before writing any of this down. A third French reading, taken at a stricter three-run bar, returned five domains, four of which appear in both of the other two. The French core is real. What moves is everything around it.


The right summary is core and rim, not “France is unstable”

The tempting sentence is that the French answer is unreliable, and it is wrong in a way that matters to anyone trying to act on it.

Every market has a core that always comes back, and the cores are close in size: 4 in the United States, 6 in Germany, 5 in Spain, 4 in France. A brand sitting in that core is being recommended consistently, in France exactly as much as in the United States.

What differs by a factor of four is the rim: 0 domains in the United States, 1 in Germany and Spain, 4 in France. The rim is the part of the answer that a single reading decides by coin flip.

So the question to ask about any AI visibility number is not “how stable is this market”. It is which of the two am I in, and one reading cannot tell you. Two can.


Market or category? The test that decides

A four-market table measured in one category is not a finding about markets. It could just as easily be a finding about CRM, and our own earlier research is the reason to take that seriously: when we measured nineteen agency buying questions across four markets, the within-market spread across verticals was as large as anything between markets, and that study concluded in print that the market is not a modifier you can apply to a number.

So before publishing this we ran two more categories in the two extreme markets, with the retirement rule written down first: if any French category rose above any American one, this was category variation and the finding was dead.

CategoryUnited StatesFrance
CRM100%50%
Accountingidentical set60%
Project management100%43%

The bands do not cross. The highest French reading is 60% and the lowest American one is 100%.

And this does not overturn the agency study, because the two measure different things. That study asked, inside a single reading, what fraction of everything cited survived every run. This one asks whether the surviving set itself comes back when you ask again. A market can cite a great many one-off sources, scoring low on the first measure, and still reproduce its durable core perfectly, scoring high on this one. Both numbers are real and they are answers to different questions.

The accounting row for the United States carries an asterisk on purpose. Its second reading came back at a one-run bar rather than two, so it cannot enter a comparison and it is not counted anywhere above. It is mentioned because of what it showed: at that different bar, it returned the same thirteen domains, name for name.


Asked at a stricter bar, and asked again later the same day

Every cell above is read at a two-run bar, which is the loosest bar we use. The obvious question is whether any of it survives asking more, so 6 of these cells were read again at three-, four- and five-run bars, each pair matched, and kept apart from the table above rather than pooled: a Jaccard between two readings is not comparable across bars, because unanimity over five is a stricter test than unanimity over two.

MarketCategoryBarComes back
United StatesCRM5 runs100%
United StatesAccounting3 runs100%
GermanyAccounting5 runs100%
SpainAccounting2 runs100%
FranceAccounting4 runs75%
GermanyCRM3 runs67%

4 of the 6 come back identical, and the lowest is 67%. So the effect is not an artefact of a loose bar. France is still the lowest market present, and the American cells are still perfect at bars where the set is twelve domains rather than four.

And one row here retires a sentence we published this morning. Germany answers 67% on CRM and 100% on accounting, both matched-bar, which is a third of the set moving inside a single market. “It is the market and not the category” was only ever sayable about the two extremes, where the category has nothing left to move: France cannot go below the American floor because the American floor is 100%. A market in the middle is where that claim can fail, and it does.

The same question, the same bar, five hours later

The American CRM string is the most reproducible cell in this study. Read at a five-run bar it returns the same twelve domains twice, and a third reading at a four-run bar returns those twelve again. Earlier the same day, the identical string at the same five-run bar returned four.

All four are inside the twelve, so nothing was lost. The answer widened, by a factor of three, in a few hours, and it is not the bar doing it: raising a unanimity bar can only remove domains from a durable set, never add them, so the later answer is citing at least twelve domains in every one of its five runs where the earlier one cited four in every one of its five.

That is the uncomfortable half of everything above. Within a session, this cell is as reliable as a measurement of this kind gets. Across an afternoon, two thirds of what it says today was not there this morning. Repeatability inside a session is not evidence that a reading taken later the same day will agree with it, and a dashboard that samples once a day is sampling a set that can triple between samples.


What to do with this

Never act on one reading. This is the whole practical content of the study and it costs one extra call to satisfy. If a source appears in two readings taken minutes apart, it is in the core. If it appears in one, you have learned almost nothing about it.

Ask a vendor how many readings its number came from, which is a different question from how many runs. A tool can average ten runs inside one reading and still show you a set that will not reproduce when it asks again tomorrow. Runs inside a reading and repeated readings are not the same control and they do not fix the same problem.

Expect a French report and an American report to differ in reliability, and do not read that as one client’s site being worse than another’s. On this evidence most of the difference is in the market, and some of it is in the question: Germany moves by a third between two categories.


What this does not show

  • One surface. This is Google’s AI Overview only. ChatGPT, Gemini and AI Mode retrieve differently and would need their own measurement, and our own work keeps finding that a result on one surface says little about another.
  • Two readings per cell. A Jaccard between two readings is one observation, not a distribution. A third reading per cell would turn this spread into a range, and that is the obvious next step rather than more markets.
  • Three categories, four markets, one day. The market comparison across categories rests on the United States and France only. Germany and Spain have one category each.
  • Minutes, then one afternoon. The tables are within-session. The one cell we read again hours later had tripled, so what is untested here is days and weeks, not merely months, and one cell is not a rate.
  • Nothing here is causal. We can say the markets differ on this measure. We cannot say what about those markets produces the difference, and we did not test any explanation.

Common Questions About Repeating a Measurement

What does “comes back” actually mean here?

It is a Jaccard index between two readings of the identical question. Each reading keeps only the domains cited on every run of that reading, so it is already a durable set rather than everything that got mentioned. The figure is the size of the overlap divided by the size of the union: 100% means the two readings named exactly the same domains, and 50% means half of what appeared across the two readings appeared in only one of them.

Why two runs per reading and not more?

Because every reading in the comparison has to sit at the same bar, and two was the bar every market could actually deliver in one session. A durable set built from five runs is tougher than one built from two, so mixing them would measure the bar rather than the market. Where a cell came back at a different bar we left it out of the comparison and said so.

Does this mean an AI visibility tool is useless in France?

No, and the core and rim split is why. A French brand in the core is cited every time we ask, exactly as consistently as an American one. What a single French reading cannot tell you is which side of that line you are on, and the fix is one more reading rather than a different tool.

Is this the same as saying AI answers are random?

It is close to the opposite. Between 3 and 6 domains come back every single time in every market we measured, including the least reproducible one. There is a stable structure in all four. The disagreement between readings lives entirely at the edges of the answer.

How would I check this on my own category?

Pick your buying question, ask it twice in the same hour with the same number of runs each time, and write down the domains cited on every run of each reading. Compare the two lists. If they match you are looking at a core; if they half-match, any single-reading report about your category is telling you less than it appears to.

Why does the accounting row for the United States not have a number?

Because its second reading answered one run where the first answered two, and a one-run reading makes every cited domain “durable” by construction. Putting it in the table would have compared two different things. It stays in the article because what it did show is worth knowing: at that different bar, it returned the identical thirteen domains.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)