Skip to content
AI VisibilityResearch
EN

We Asked Four AI Surfaces the Same Question. Fifty-One Sources Came Back and Not One Appeared on All Four.

Same question, same market, same moment, four surfaces. Forty-one of fifty-one cited sources appeared on exactly one of them. Google's own three barely overlap.

· Updated · 10 min read

Every study we have published carried the same limitation: one surface, Google’s AI Overview. This one tests what that limitation costs.

We put two questions to four surfaces at the same moment: AI Overview, Google AI Mode, Gemini and ChatGPT. Between them they cited 51 distinct domains. Not one appeared on all four. 41 appeared on exactly one.

The uncomfortable part for Google is that three of those four surfaces are theirs, and AI Overview and Gemini did not share a single source on either question.

How this was measured: two prompts, United States, English, 7 August 2026, one run per surface, all four surfaces queried in the same call. One run is a draw and not a rate, which is why this article counts overlap between surfaces rather than frequency within them. Cited domains only, never brand mentions. Both runs are in our measurement register with the per-surface counts.


Question one: running shoes

SurfaceDomains cited
ChatGPT11
AI Overview5
AI Mode4
Gemini3

20 distinct domains. One appeared on three surfaces (runrepeat.com), one on two, and 18 on exactly one.

Question two: ecommerce email platforms

SurfaceDomains cited
ChatGPT16
AI Mode11
AI Overview8
Gemini5

31 distinct domains. One appeared on three (drip.com), seven on two, and 23 on exactly one.

The two questions together

Domains appearing on…CountShare
All four surfaces00%
Three surfaces24%
Two surfaces816%
Exactly one surface4180%

Four fifths of the cited web, for these questions, is visible to exactly one of the four places a buyer might ask.

All four surfaces0 (0%)Three surfaces2 (4%)Two surfaces8 (16%)Exactly one surface41 (80%)
Fifty-one distinct domains across two buying questions. Not one appeared on all four surfaces, and four fifths appeared on exactly one. US English, AI Overview, AI Mode, ChatGPT and Gemini, one run per surface, which is a draw and not a rate.
How many of the 51 cited domains appeared on one, two, three or all four AI surfaces
value
All four surfaces0
Three surfaces2
Two surfaces8
Exactly one surface41

Why this is worse than “engines differ”

Everyone in this category will tell you engines differ. What the numbers add is how much, and one specific result nobody advertises.

AI Overview and Gemini cited nothing in common. On either question. Both are Google. Ask Google’s search AI and Google’s assistant the same thing in the same minute and the evidence underneath the two answers is disjoint.

AI Mode, also Google, shares one domain with AI Overview on the first question and four on the second. Better, and still mostly separate.

If you are measuring one Google surface and calling it “Google”, you are measuring one Google surface.

The result this sits next to

Here is what makes the finding sharp rather than merely interesting.

Running shoes is the most stable category we have ever measured. In our stability study, three runs of that exact question on AI Overview returned five domains and all five appeared in every single run. Perfect within-surface reproducibility.

That same question, asked of four surfaces, produced twenty domains of which eighteen appeared on exactly one.

Reproducibility within a surface tells you nothing about agreement between surfaces. A vendor can honestly report that your measurement is rock solid, repeat it all month, and still be describing one quarter of the picture. Those are two different properties and this market routinely sells the first while implying the second.

ChatGPT cites the most, Gemini the least

ChatGPT returned the largest source list on both questions, 11 and 16. Gemini returned the smallest, 3 and 5. AI Overview and AI Mode sat between them.

We are not going to tell you what that means about answer quality, because we did not measure quality. What it means practically is that the surface you measure changes not only which sources you see but how many, and a share-of-voice percentage computed against a three-domain denominator is not comparable with one computed against sixteen.

That is a straightforward arithmetic problem hiding inside a lot of dashboards.

What this changes about buying

A per-engine price is a per-picture price. Our verified catalogue shows entry plans covering one engine (Surfer, and our own €29 plan), three (Siftly, Writesonic), four (Nightwatch) and six (Promptmonitor at $29, Rank Prompt). On this evidence the difference between one engine and four is not a refinement. It is most of the data.

Do not average across engines. With 80% of sources appearing on a single surface, a blended visibility score mixes four largely disjoint populations into one number that describes none of them. Read them separately or do not read them.

Check which surface your buyers actually use before optimising for the one your tool happens to cover. That is a question about your market, not about AI.

And remember the other axis. How many runs a category needs varies from about three to about thirty, and this article adds a second dimension on top of it. Coverage and repetition are separate purchases and most vendors sell you a fixed amount of both.

Replicated five more times on 7 August 2026

This article measured two consumer questions. Since writing it we have run the same instrument five more times, on questions with nothing in common with running shoes. The headline held for four of them and broke on the fifth.

QuestionMarketCited domainsOn two surfacesOn three or four
Best GEO agencyUnited States2050
Best GEO agencyFrance1940
Best GEO agencySpain1640
Which AI visibility toolUnited States2110
Best GEO agencyGermany, added 7 Aug 20261822

Different category, four languages, and one of the runs is business software rather than a service. In the first four replications, not one source was cited by three of the four surfaces. Germany, run last, broke that. Two domains reached three surfaces there, and both are agency websites: one belongs to Aufgesang and one to Suchhelden. Four runs made it look like a rule and the fifth says it is not one, which is the whole reason to keep replicating. On the tools question Google AI Overview returned a server error, so that row is three surfaces rather than four, and it is our instrument failing rather than Google going quiet.

Add the two questions in this article and that is seven question sets across two categories and four languages. Not one produced a source cited by all four surfaces, and that is the part that has held every time. On three surfaces the picture is thinner but not empty: the two consumer questions here managed exactly one source each, four of the five newer runs managed none, and Germany managed two. We are not calling any of it a law, because every one of them is a single pass and we know what a single pass is worth. What we can say is that the all-four column has never moved off zero in seven tries, and that the three-surface column moves depending on whether somebody in that market has published a page worth agreeing on.

What this does not show

One of the four surfaces returns less than it cites, and that reaches the headline. Measured on 15 August 2026: an AI Mode answer prints numbered citation markers in its body and the structured list we read carries about half of them, while AI Overview’s carries all of them. So “not one appeared on all four” is partly a statement about AI Mode’s capture: a domain present on every surface could be missing from the AI Mode column and break the run without ever being absent from the answer.

What survives untouched is the sharper half. AI Overview and Gemini cited zero domains in common on both questions, and both of those lists carry every marker their answers print, so that pair is measured with one instrument on both sides. ChatGPT’s capture we cannot check at all, because that call times out on this route.

  • One run per surface. Two questions, eight surface-answers. This measures how little the four source sets overlap at one moment. It does not measure how often any given domain appears, and it cannot: for that you need repetition, which is a different and more expensive design.
  • Two questions, one market, one language, one day. Both questions are consumer or SMB software in the United States in English.
  • Overlap is not quality. Four surfaces disagreeing does not make any of them wrong. Different evidence can support the same recommendation, and in fact all four named overlapping products while citing different sources.
  • We counted domains, not prominence. A domain cited once and a domain the answer leaned on are counted the same.
  • Nothing causal, and nothing about brand mentions. This is citation data only.
  • Surfaces move. These are four products under active development, measured on one afternoon.

Common Questions About Differences Between AI Engines

Do ChatGPT, Gemini and Google AI Overview cite the same sources?

Barely. Across two questions asked of four surfaces at the same moment, 51 distinct domains were cited and not one appeared on all four. 41 of the 51, about 80%, appeared on exactly one surface.

Do Google’s own AI surfaces agree with each other?

Not on our evidence. AI Overview and Gemini cited zero domains in common on both questions. AI Mode, also Google, shared one domain with AI Overview on the first question and four on the second. Measuring one Google surface does not tell you about the others.

Which AI engine cites the most sources?

ChatGPT, on both questions we tested, with 11 and 16 distinct domains. Gemini cited the fewest, with 3 and 5. That matters for any percentage: a share-of-voice figure calculated against five sources is not comparable with one calculated against sixteen.

Is it worth paying for a tool that covers more engines?

On this evidence the coverage difference is most of the data rather than a refinement, since four fifths of cited sources appeared on only one surface. Entry plans in our verified catalogue range from one engine to six at similar prices, so the comparison is worth making deliberately.

If my measurement is stable, am I seeing the whole picture?

No, and this is the trap. The question in this study with perfect within-surface stability, five domains cited in every one of three runs, produced eighteen single-surface domains out of twenty when we asked four surfaces. Reproducibility and coverage are different properties, and a tool can be excellent at the first while silent about the second.

Can I reproduce this?

Yes. Two prompts, United States, English, four surfaces, one run each, on 7 August 2026, with both prompts quoted in full above and the per-surface counts in our register. Ask the same question of four surfaces in the same session and list the domains each one cites.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)