The Same Question in Three Markets, Three Different Answers
We asked which tool tracks ChatGPT visibility in the US, Spain and France. Not one tool was recommended consistently in more than one market.
We asked Google’s AI Overview the same question in three markets: which tool should I use to track my brand’s visibility in ChatGPT? In the United States in English, in Spain in Spanish, in France in French.
Twelve runs. The three markets cited 16, 17 and 13 domains, and they barely overlap. Not one tool was recommended consistently in more than one market.
The French answer named a freelance SEO consultant in Nice, in every single run, ahead of every funded platform in the category.
Disclosure: EchoWi is our product and competes in this category. That is a conflict, and the answer to a conflict is a method you can check rather than a promise to be neutral: the question, the surface, the markets, the run counts and the dates are all stated above, so any of it can be repeated against us.
The short version
- No tool appears stably in more than one market. The stable sets do not overlap at all.
- Local sites dominate locally. France’s answer is French sites. Spain’s is Spanish sites. The United States gets American ones.
- Not one funded vendor appeared, anywhere. Zero across twelve runs and three markets for Profound, AirOps, Scrunch, Peec AI, Evertune and Semrush.
- In France, a solo consultant beat the entire category, cited in every run.
- Only one domain came close to crossing borders:
seranking.com, stable in two markets and at 67% in the third.
Method
| Question | The same in each language: which tool is best for tracking a brand’s visibility in ChatGPT |
|---|---|
| Surface | Google AI Overview |
| Markets | United States and English, Spain and Spanish, France and French, all set explicitly |
| Runs | 6 in the US, 3 in Spain, 3 in France |
| Cache | Bypassed. Every run is a fresh upstream call |
| Execution | Serial |
| Dates | 5 and 6 August 2026 |
Setting the market explicitly is the whole experiment. Measurement tools default to the United States and English. If you leave that alone, a question in French gets answered by the American market and nothing in the report tells you. That would have produced three identical results and a completely wrong conclusion.
The limitation, stated plainly: the question asks about ChatGPT, but the surface measured is Google’s AI Overview. It is one of the ways people research this question, not the only one. The US sample is twice the size of the other two.
The result, side by side
| United States | Spain | France | |
|---|---|---|---|
| Runs | 6 | 3 | 3 |
| Domains cited | 16 | 17 | 13 |
| Cited in every run | 5 | 3 | 4 |
| Cited exactly once | 8 | 12 | 7 |
| cited in every run | distinct domains cited | |
|---|---|---|
| United States, 6 runs | 5 | 16 |
| Spain, 3 runs | 3 | 17 |
| France, 3 runs | 4 | 13 |
And the domains cited in every run, which is where the finding lives:
| United States | Spain | France |
|---|---|---|
| reddit.com | useomnia.com | invox.fr |
| seranking.com | surfeo.ai | redback-optimisation.fr |
| youtube.com | blog.hubspot.es | seranking.com |
| siftly.ai | youtube.com | |
| alhena.ai |
Look at the tools specifically. The United States gets siftly.ai and alhena.ai. Spain gets useomnia.com and surfeo.ai. France gets neither pair. There is no tool in common between any two of these three lists.
The only domain approaching consistency is seranking.com, an SEO suite rather than an AI visibility tool: stable in the United States and France, and at 67% in Spain.
We ran it again two days later, and the shape held while the names did not
On 7 August 2026 we re-measured all three markets. The aggregate barely moved:
| Published, 5 to 6 August | Re-measured, 7 August | |
|---|---|---|
| United States | 6 runs, 16 domains, 5 stable | 4 runs, 23 domains, 3 stable |
| Spain | 3 runs, 17 domains, 3 stable | 3 runs, 17 domains, 3 stable |
| France | 3 runs, 13 domains, 4 stable | 3 runs, 13 domains, 3 stable |
Spain returned the identical count twice: seventeen domains, three of them in every run. France returned thirteen domains both times. If you only ever looked at the totals, you would conclude this measurement is highly repeatable.
Now look at which domains those were.
| Stable on 6 August | Stable on 7 August | |
|---|---|---|
| Spain | useomnia.com, surfeo.ai, blog.hubspot.es | acumbamail.com, blog.hubspot.es, youtube.com |
| France | invox.fr, redback-optimisation.fr, seranking.com, youtube.com | seranking.com, invox.fr, meltwater.com |
One of Spain’s three survived. Both Spanish tools dropped out, and the replacement set contains no AI visibility tool at all: an email marketing company, a HubSpot blog and YouTube.
We also ran France a second time the same day, in a separate session. It produced a third answer again: seranking.com and dageno.ai stable, meltwater.com absent, redback-optimisation.fr in one run of two. Across all three French sessions now on record, exactly one domain is stable in all of them, seranking.com.
This is the useful finding, and it is not the one we published first. The size and shape of the answer are reproducible. The membership is not. A vendor reporting that you were cited in 100% of runs is describing that session, and our own numbers say the next session will disagree with it.
What we got wrong methodologically: the study above described its questions but never published the exact strings. So this is a re-measurement of the same question, not a strict replication, and we have since shown that rewording a question changes almost every source it returns. Every measurement we publish now records the prompt verbatim, in
src/data/measurements.ts, with its market, surface, date and run count.
Two things below also need marking against this. siftly.ai, stable in the United States, appeared in two of three French runs on 7 August, which is the closest any tool has come to crossing a border. And tryprofound.com appeared once in France, so the claim further down that no funded vendor appeared anywhere holds for the original twelve runs and no longer holds in general.
France recommended a freelancer
This is the single most striking result in the study, and it is also the one that did not survive re-measurement. Read it with the section above.
redback-optimisation.fr was cited in all three French runs. It belongs to Florian Zorgnotti, a freelance SEO consultant based in Nice, offering audits and ongoing optimization to French-speaking businesses.
One person, in Nice, cited in every run, ahead of every company in this category with a Series A.
And in five further French runs on 7 August, that domain was cited in one. Zero in the three-run session, one in the two-run session. The original observation was real and we are not withdrawing it; three runs simply cannot tell “always” from “sometimes”, which is the whole argument of the section above. What we can still say is the weaker and more durable version: an independent consultant was cited by this surface at all, repeatedly, in a category where funded platforms were not.
The other French domain cited in every run, invox.fr, returned a 403 to our request, so we cannot describe what it is. We are listing it because it was cited, and not describing it because we could not check it.
Nobody funded appeared, in any market
Across twelve runs in three markets, the following appeared zero times:
| Vendor | Raised or valued at |
|---|---|
| Profound | $155M+ raised, $1B valuation |
| AirOps | ~$118M raised |
| Scrunch | Acquired for a reported $225M |
| Peec AI | $29.1M raised |
| Evertune | $19M raised |
| Semrush | Public company |
We ran this expecting the funded vendors to dominate at least their home market. They did not appear in it at all.
What this means
-
A comparison article written in one market does not describe another. If you are in France choosing a tool, an English-language “best GEO tools” list is describing a different answer to your question. The tools your own market’s AI recommends are not on it.
-
Language is not a translation problem, it is a retrieval problem. This matches what has been published: in Temso’s analysis of over 7 million citations, French prompts were answered with French-language sources 82.1% of the time in AI Overviews. Our three markets are a much smaller sample pointing the same way.
-
Scale is not the advantage people assume. A consultant in Nice and two tools nobody in the industry discusses beat every funded platform. Retrieval does not know what anyone’s valuation is.
-
If you sell in several markets, you have several visibility problems. Not one problem measured from head office. That is the practical, expensive implication, and it is the reason we set country and language explicitly on every measurement we publish.
The mistake that would have hidden all of this
This study exists because of one setting, and it is worth being concrete about it, because getting it wrong produces a result that looks perfectly healthy.
Almost every measurement tool in this category defaults to the United States and English. That default is not announced in the report. Ask a Spanish question without changing it and the answer comes back from the American market, formatted identically to a Spanish result, with nothing on screen to mark the difference.
Had we left the default alone, all three of our runs would have returned roughly the same American domains, and the honest conclusion would have been the opposite of the real one: that AI recommends the same tools everywhere, so one measurement covers you. That conclusion would have been comfortable, cheap and wrong.
What to check in any report you are shown, whether ours or a vendor’s:
- Which country and language were set. If the report does not say, assume the default.
- Whether each market was run separately. An average across four markets can hide being invisible in three of them.
- How many runs sit behind the number. Answers vary between identical requests, so one run is a draw rather than a rate.
- Whether the cache was bypassed. A cached answer repeated five times is one measurement shown five times.
- Which domains were cited, not only whether you were named. That is the part you can act on.
None of those is expensive to ask for, and a tool that cannot answer them is producing a number nobody can defend.
Since publishing this, we found the same shape twice more
The obvious objection to the study above is that GEO software is a strange, tiny category and nothing here generalises. So we ran the same design in two sectors that have nothing to do with it.
Flights, and a correction. This section originally reported that the United States flight answer cited 6 domains with only Reddit appearing every time, and that no booking site was cited more than once in five runs. Re-measured with the corrected instrument on 7 August 2026, the same question returned 6 cited domains and all six appeared in every run. The instability was not there. It came from a script that counted the searches behind an answer alongside the sources the answer actually cited, which inflated the domain list with one-off entries.
The corrected figure is registered and the full cross-category comparison is in our study of how stable AI answers are by category, where flight booking sits with the most stable categories we have measured rather than the least.
Universities. We asked which are the best private universities in Madrid, five runs, and the best universities in Paris, three runs. Both named the institution we were tracking every time. Neither cited a single university website in its stable set. The sources answering both questions were student residences, a language school, a relocation agency and education directories.
Three sectors, four countries, three languages, and the shape holds:
| Stable citations went to | |
|---|---|
| GEO software, US | Reddit, YouTube, two tools nobody discusses |
| GEO software, France | An independent consultant in Nice |
| Flight booking, Spain | Two metasearch engines, Reddit, Instagram, TikTok |
| Flight booking, US | Six domains, every one of them in every run |
| Universities, Madrid | Two student residences and a guidance platform |
| Universities, Paris | A language school, a relocation agency, a tutoring firm |
The pattern is not that big brands lose. It is that the stable citations go to whoever published content that compares, explains or discusses, and that is rarely the company whose product is being compared. A booking form, a syllabus and a vendor product page have nothing for a retrieval layer to extract, whatever the brand on them.
A second pattern used to sit here, and it was wrong. We reported that the American answer was less stable than the European one in both categories, and built a hypothesis on it about contested markets leaving retrieval without a decisive answer. That pattern was an artifact of the counting error described above. Re-measured, the United States flight answer is completely stable: six domains, all six in every run.
The hypothesis survived in a different form, which we since tested properly across six categories. Instability does not track the market. It tracks the category: settled product categories return the same sources every time, while professional services and emerging software categories return a different subset on almost every run. That is a separate study with its own data, and it is the version we would now defend.
What this does not show
- Nothing about product quality. Citation is not a review. No controlled accuracy comparison of tools in this category has been published by anyone, us included.
- Nothing causal. A July 2026 review of 45 GEO studies found no technique with a demonstrated causal, stable, cross-platform effect. Twelve runs do not change that.
- Nothing about other surfaces. This is Google’s AI Overview. ChatGPT, Perplexity and Gemini would each need their own measurement.
- Nothing durable. These are answers on three days in August 2026, and the re-measurement above shows the membership moving between two of them. Repeat it in a month and expect movement.
Common Questions About This Study
Which tools does Google’s AI recommend for tracking ChatGPT visibility?
It depends entirely on the market, and on the day. On 6 August, siftly.ai and alhena.ai were cited in every US run, useomnia.com and surfeo.ai in every Spanish one, and neither pair in France. On 7 August the Spanish stable set contained no tool at all. Treat any answer to this question as a photograph, including this one.
Why do the answers differ so much between countries?
Retrieval pulls strongly toward sources in the language of the question, which is documented across engines and languages. A French question is answered largely from the French-language web, where a different set of pages exists.
Is twelve runs enough to conclude this?
It is enough to establish that the stable sets do not overlap, which is the claim being made. It is not enough for precise rates, and the US sample is twice the size of the others. Every run is billed and executed serially, so sample size has a real cost.
Re-measuring sharpened this. Twelve runs was enough for the aggregate, which reproduced almost exactly, and not enough for any statement about a specific domain: three runs called a source “cited in every run” that five later runs cited once. Runs within one session tell you about that session. To claim a domain is durably cited, you need sessions on different days, which is a different and more expensive design than anyone in this category currently sells.
Why did no funded vendor appear?
The measurement cannot say for certain. The likeliest explanation, and the one best supported by published research, is that retrieval favours pages answering the question directly, and vendor marketing pages are written to sell rather than to answer. Funding buys sales capacity, which a retrieval layer cannot see.
Should I choose a tool based on what the AI recommends?
No. Two of the five stable US domains were not products at all, and the French answer named a consultant rather than software. Use it to understand what your buyers are being told, not as a shortlist.
Can I reproduce this?
Yes. The question, surface, markets, languages, run counts and dates are all above. If you get a different answer, it is because the model changed or because you did not set the market explicitly.
Where this leaves you
The result we did not expect was the absence, not the presence. Six companies holding more than $200 million between them appeared zero times, across twelve runs and three markets, on the query that describes their own product.
And on the day we measured, the winner in France was one person with a laptop in Nice.
If you are trying to be recommended by these systems, that is either the most encouraging finding in this article or the most alarming one, depending on which side of it you are standing. Re-measuring added a third reading, which is the one we would actually act on: the position is winnable and it is not holdable. A source that owns a question this week is not the source that owns it next week, and the practical answer to that is to measure on a schedule rather than to celebrate a screenshot.
We since ran this design across six categories to find out whether the churn is a property of AI answers or a property of this category. It is the category: some categories cite the same sources every single time, and ours is not one of them.
Next: the full US measurement, and what GEO is and what the research supports.
See what 12 AI engines say, per country →
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.