Skip to content
AI VisibilityGEO
EN

Half of What an AI Cites Is Gone in Six Days. The Other Half Barely Moves.

Fourteen buying questions asked twice in August and again six days later. The full cited set overlaps 50 per cent. The durable core holds 73 per cent.

· Updated · 15 min read

Every AI visibility tool wants to sell you a chart of your most-cited sources over time. We built one from real measurements and then asked the question the chart cannot answer: when a source drops off it, did it actually leave?

14 buying questions, asked on 15 and 16 August and again on 22 August. Same surface, same market, same language, same number of runs every time. The whole cited set overlaps 50 per cent between the baseline and six days later. The set of domains that came back in both baseline readings holds at 73 per cent.

So half of what an assistant cites is different within a week, and the durable half is a different thing that behaves differently. A chart that reports one number cannot tell you which half moved.

Transparency and method: EchoWi sells AI visibility measurement, so a study about whether these charts mean anything is a study about our own category. The design, the deciding statistic, both predictions and the retirement threshold were written down and committed before the first call. The questions, the surface, the market, the dates and the number of runs answered per cell are all in our measurement register, and with that anybody can run it against us.

The short version

  1. The whole cited set turns over about half in six days. Median overlap 50 per cent across 12 comparable cells.
  2. The core holds at 73 per cent. Core means the domain came back in both baseline readings rather than appearing once.
  3. Durability is a property of the question, not of the method. 4 cells kept every core domain. 4 kept exactly half. Same surface, same market, same bar, same six days.
  4. Three cells that looked perfectly stable were not. All three gained churn once a third reading existed.
  5. Rephrasing the question costs almost nothing. 5 of 6 intents returned an identical set from a keyword and a question form on the same day.

Why a share is the wrong unit

Before the time question, there is a measurement question, and getting it wrong invents the answer.

Ask the same thing twice and the number of sources cited once moves a lot. The number cited every time barely moves. So a percentage built as “durable over total” has a nearly constant numerator and a denominator that wanders, which produces a figure that looks like a measurement and behaves like a coin toss.

There is a second trap underneath it, and it is worth more than most people would guess. More runs answered means more chances for a domain to appear at all, so a reading with ten runs has a bigger set than one with three. Compare a ten-run reading against a three-run baseline and the core will look like it collapsed, when what actually happened is that the second reading had more opportunities to add strangers.

We measured that confounder directly across the cells in our register with two or more dates:

CellsCoreBorderCore share
Mixed number of runs10145418 %
Same number of runs668543 %

The 18 per cent was measuring the bar and not the world. Twenty five points of difference, from a variable nobody puts on the chart. Every figure in this study uses cells that answered the same number of runs on both dates, and the two cells that came back with fewer are excluded and counted separately rather than compared at a lower bar.

Six days later

12 cells qualified. Retention is the fraction of the baseline core still cited on 22 August.

QuestionCoreStill citedRetention
best grammar checker44100 %
best note taking app22100 %
what is the best payroll service for a small business?5480 %
best crm for small business4375 %
what is the best AI coding assistant?10770 %
best vpn service3267 %
best payroll service for a small business10550 %
best website to book flights2150 %
what are the best running shoes for beginners?4250 %

Median 73 per cent, against a median overlap of 50 per cent for the whole cited set. The core is the durable half by 23 points.

And the distribution matters more than the median. 4 cells kept everything and 4 kept exactly half. There is not much in between. Whether your category has a stable answer is a property of your category, and no amount of sampling turns an unstable one into a stable one.

The cells that looked perfect were the ones with nothing to get wrong

3 cells returned a baseline where every cited domain appeared in both readings. Zero churn. Read quickly, that is the dream: a fixed set of sources you can go and work on.

We predicted before measuring that at least 2 of the 3 would gain churn with a third reading, because a set that agrees across two readings has passed a very low bar rather than proved itself fixed.

All 3 did. One went from ten domains with no churn to seven domains sharing five with the baseline. Another went from ten to eleven sharing seven.

That is the part a weekly chart hides. A source that has appeared twice is not established, and a tool that shows you two data points and a line between them is showing you a line.

Rephrasing costs almost nothing, and the exception is the same one twice

Six of the fourteen cells are the same buying intent in two forms, a bare keyword and a full question. On the same day, at the same bar:

5 of 6 returned an identical set of domains. Not similar, identical.

The sixth is payroll, and it diverged in the baseline too. Two readings a week apart, and the same intent is the one that splits both times while the other five hold. That is not noise, that is a property of that question.

For anyone buying a tracker this is the useful half. The worry that your numbers depend on exactly how the prompt was worded is real, and it is mostly concentrated in a minority of intents rather than spread evenly. But you cannot know which of your intents is the unstable one without asking it more than one way, and almost nothing on the market does that.

Why the churn half exists at all

An assistant re-plans retrieval on every request. It is not reading a fixed index of trusted sources and quoting from it, and the sources that survive every reading of a question look different from the ones that appear once. In the cells here the durable half is dominated by a small number of large destinations that show up across many unrelated questions, while the churn half is long, thin and mostly made of pages that fit one phrasing of one question on one day.

That is why the two halves need reporting separately rather than summing. Working on the durable half is a different job from working on the churn half, and only one of them is a job you can plan.

What this means if you are buying AI visibility measurement

The chart is not useless. It is reporting two different things added together, and the useful one is the smaller one.

  • Ask what the number counts. A percentage of durable sources over total sources moves mostly with the total. Ask for the set.
  • Ask how many runs, and whether that number is fixed. A comparison across weeks where the number of answered runs drifts is measuring the drift. We measured that at twenty five points.
  • Ask whether the tool runs more than one wording per intent. 5 of 6 intents do not care, and finding the sixth is the whole value.
  • And ask whether it separates the core from the churn, because six days moved half the set and left three quarters of the core alone. Those are two different stories about your visibility and only one of them is worth acting on.

If you want that separation on your own category rather than on ours, that is what our product does, and the method above is the one it uses.

The surface is not the bigger axis, and we predicted it would be

This study’s stated limit was one surface, so we closed it: the same 12 cells read on AI Mode the same day, so the pairing is inside the cell and only the surface changes.

The prediction, written down and committed before the first call, was that the surface would matter more than a week. It does not.

Median overlap
Same surface, six days apart50 %
Both surfaces, same day50 %

That is the retirement threshold the design named, so the claim that mixing surfaces adds more noise than a month of drift is retracted here rather than quietly dropped.

What the Jaccard was hiding is better than what we predicted. A low overlap between a wide set and a narrow one can describe containment rather than disagreement, and this is containment: a median 94 per cent of the AI Overview set sits inside the AI Mode set, which is 1.52 times wider. 6 of the 12 cells are fully contained, and only 2 fall below half.

So the two Google surfaces are not answering with different sources. One is answering with a superset of the other, and the Jaccard punishes the narrow side for being narrow. If you are cited in AI Overview you are almost certainly cited in AI Mode, and the reverse is a coin toss.

The two cells that break it are worth naming because they are real disagreement rather than width: both phrasings of the flights question share only 3 of 7, with AI Mode answering from Skyscanner, Frommers and US News where AI Overview answered from Kayak, Expedia and Quora. Same day, same market, same question, two different travel sections.

Two Google surfaces nest. The third does not.

Containment is the finding above, and it came from two surfaces that share an index. That leaves a question a buyer needs answered: does one surface contain another because that is how retrieval works, or because those two are the same retrieval seen at two widths?

Gemini on the same 12 cells, same day, separates them. It does not nest.

Same 12 cells, same dayMedian containment
AI Overview inside AI Mode94 %
Gemini inside AI Mode42 %
Gemini inside AI Overview37 %

That was the retirement threshold the design named, so the claim that a narrow set sits inside a wide one is now bounded to the pair it was measured on, and not published as a property of surfaces.

And the pair of numbers is the whole point. Gemini against AI Overview has a Jaccard of 0.25, and AI Overview against AI Mode has 0.50. Those look like the same kind of number, twice as much agreement in one case. They are not the same kind of number at all. One is a narrow set living inside a wide one; the other is two sets that genuinely do not share sources. Only the containment figure separates them, and no tool on the market reports it.

Gemini is also slightly narrower than AI Overview, a median 0.87 of its size, so this is not a width story. It is a different retrieval wearing the same company’s name.

What that means if you are spending money on measurement: watching AI Overview tells you most of what you would learn from AI Mode, and close to nothing about Gemini. Five of the eight comparable cells share less than half of Gemini’s sources with AI Overview. If your tool reports one Google number, ask which surface it came from, because the other one is a different market.

And the nesting is American

If one Google surface contains the other because they share an index, that has to show up in any market. So the same eight intents in German, in Germany, on both surfaces, in one call each so the two readings cannot drift apart.

It shows up weaker.

Median containment, AI Overview inside AI Mode
United States94 %
Germany, all comparable cells50 %
Germany, the 5 cells read at three runs67 %

50 per cent was the retirement threshold the design named, so the American figure is now a statement about that market and not about Google’s architecture.

The two numbers for Germany are both honest and they differ for a reason worth saying out loud. Containment is a fraction of the AI Overview set, so when both surfaces answer fewer runs both sets get narrower and the fraction falls. The two cells that answered twice are the two lowest in the table. That biases this arm downward, which is the opposite of the Gemini arm, where only one side’s bar moved and a low bar pushed containment up. A run count is not a neutral setting; which direction it bends a figure depends on which side of the fraction it narrows.

And the control is what makes the arm mean anything. If the German answers had used the same sources as the American ones, this would be one market measured twice. They do not: joined question by question, the median overlap between the German and American source sets is 0.11, running from 0.03 to 0.27, and the two lowest cells share a single domain out of thirty odd. Germany is a different set of sources, and inside it the two surfaces relate differently.

One German cell returned no AI Overview across three runs. That is a surface declining to appear, not a zero, and it is excluded and named rather than counted.

What this costs a buyer: the shortcut is local. In the United States, watching AI Overview tells you most of AI Mode. In Germany it tells you between half and two thirds, and the German AI Mode set is 1.67 times wider. If you sell in more than one country, the two surfaces are two measurements there, not one.

What this does not show

  • Three Google surfaces and two markets. Retention is AI Overview in the United States only; the surface halves add AI Mode and Gemini on the same cells, and the market half repeats the containment measure in Germany. ChatGPT and Perplexity retrieve differently and by this route ChatGPT returns no sources at all, so its zero would be the instrument.
  • Two markets and two languages. Everything before the market section is the United States in English; the containment measure is repeated in Germany in German, and the cross-market control is in the register.
  • Two points and then a third. Six days is one interval, not a trend, and nothing here says the churn rate is constant.
  • 12 comparable cells. A median over 12 is a direction and not a precise figure, and the two cells excluded for answering at a lower bar are named in the register rather than quietly dropped.
  • Nothing here is causal. We measured what changed, not why.

Frequently asked questions

What is the difference between the core and the rest?

The core is the set of domains cited in every reading of a question rather than in some of them. The rest appeared at least once. The distinction matters because the two behave differently over time: in this study the core retained 73 per cent over six days and the full set overlapped 50 per cent.

Does asking more times make the number more stable?

Not in the way people expect. More runs raises the number of sources seen at least once, so a strict “cited every time” statistic gets harder to satisfy as you sample more. That is why comparisons across time have to hold the number of runs fixed, and why we measured a twenty five point difference between mixed and matched bars.

Is a source that disappeared from the chart actually gone?

Often not. In this study, half of the full cited set differed after six days while three quarters of the core stayed. A domain that drops out of a weekly view may simply have been in the churn half the whole time, which is why the useful chart shows the core and the border separately instead of one line.

How many times should a tool ask each question?

Enough to tell a stable citation from a coincidence, and the same number every time it compares. What matters more than the exact figure is that the number does not drift between the weeks being compared, because then the comparison measures the sampling and not the market.

Does the wording of the prompt change the answer?

For most intents, no. 5 of 6 intents here returned an identical set from a keyword form and a question form on the same day. The sixth returned different sets on both dates it was measured, which is why the honest approach is to ask each intent more than one way rather than to assume the wording is safe.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)