Half of What an AI Cites Is Gone in Six Days. The Other Half Barely Moves.
Fourteen buying questions asked twice in August and again six days later. The full cited set overlaps 50 per cent. The durable core holds 73 per cent.
Every AI visibility tool wants to sell you a chart of your most-cited sources over time. We built one from real measurements and then asked the question the chart cannot answer: when a source drops off it, did it actually leave?
14 buying questions, asked on 15 and 16 August and again on 22 August. Same surface, same market, same language, same number of runs every time. The whole cited set overlaps 50 per cent between the baseline and six days later. The set of domains that came back in both baseline readings holds at 73 per cent.
So half of what an assistant cites is different within a week, and the durable half is a different thing that behaves differently. A chart that reports one number cannot tell you which half moved.
Transparency and method: EchoWi sells AI visibility measurement, so a study about whether these charts mean anything is a study about our own category. The design, the deciding statistic, both predictions and the retirement threshold were written down and committed before the first call. The questions, the surface, the market, the dates and the number of runs answered per cell are all in our measurement register, and with that anybody can run it against us.
The short version
- The whole cited set turns over about half in six days. Median overlap 50 per cent across 12 comparable cells.
- The core holds at 73 per cent. Core means the domain came back in both baseline readings rather than appearing once.
- Durability is a property of the question, not of the method. 4 cells kept every core domain. 4 kept exactly half. Same surface, same market, same bar, same six days.
- Three cells that looked perfectly stable were not. All three gained churn once a third reading existed.
- Rephrasing the question costs almost nothing. 5 of 6 intents returned an identical set from a keyword and a question form on the same day.
Why a share is the wrong unit
Before the time question, there is a measurement question, and getting it wrong invents the answer.
Ask the same thing twice and the number of sources cited once moves a lot. The number cited every time barely moves. So a percentage built as “durable over total” has a nearly constant numerator and a denominator that wanders, which produces a figure that looks like a measurement and behaves like a coin toss.
There is a second trap underneath it, and it is worth more than most people would guess. More runs answered means more chances for a domain to appear at all, so a reading with ten runs has a bigger set than one with three. Compare a ten-run reading against a three-run baseline and the core will look like it collapsed, when what actually happened is that the second reading had more opportunities to add strangers.
We measured that confounder directly across the cells in our register with two or more dates:
| Cells | Core | Border | Core share |
|---|---|---|---|
| Mixed number of runs | 101 | 454 | 18 % |
| Same number of runs | 66 | 85 | 43 % |
The 18 per cent was measuring the bar and not the world. Twenty five points of difference, from a variable nobody puts on the chart. Every figure in this study uses cells that answered the same number of runs on both dates, and the two cells that came back with fewer are excluded and counted separately rather than compared at a lower bar.
Six days later
12 cells qualified. Retention is the fraction of the baseline core still cited on 22 August.
| Question | Core | Still cited | Retention |
|---|---|---|---|
best grammar checker | 4 | 4 | 100 % |
best note taking app | 2 | 2 | 100 % |
what is the best payroll service for a small business? | 5 | 4 | 80 % |
best crm for small business | 4 | 3 | 75 % |
what is the best AI coding assistant? | 10 | 7 | 70 % |
best vpn service | 3 | 2 | 67 % |
best payroll service for a small business | 10 | 5 | 50 % |
best website to book flights | 2 | 1 | 50 % |
what are the best running shoes for beginners? | 4 | 2 | 50 % |
Median 73 per cent, against a median overlap of 50 per cent for the whole cited set. The core is the durable half by 23 points.
And the distribution matters more than the median. 4 cells kept everything and 4 kept exactly half. There is not much in between. Whether your category has a stable answer is a property of your category, and no amount of sampling turns an unstable one into a stable one.
The cells that looked perfect were the ones with nothing to get wrong
3 cells returned a baseline where every cited domain appeared in both readings. Zero churn. Read quickly, that is the dream: a fixed set of sources you can go and work on.
We predicted before measuring that at least 2 of the 3 would gain churn with a third reading, because a set that agrees across two readings has passed a very low bar rather than proved itself fixed.
All 3 did. One went from ten domains with no churn to seven domains sharing five with the baseline. Another went from ten to eleven sharing seven.
That is the part a weekly chart hides. A source that has appeared twice is not established, and a tool that shows you two data points and a line between them is showing you a line.
Rephrasing costs almost nothing, and the exception is the same one twice
Six of the fourteen cells are the same buying intent in two forms, a bare keyword and a full question. On the same day, at the same bar:
5 of 6 returned an identical set of domains. Not similar, identical.
The sixth is payroll, and it diverged in the baseline too. Two readings a week apart, and the same intent is the one that splits both times while the other five hold. That is not noise, that is a property of that question.
For anyone buying a tracker this is the useful half. The worry that your numbers depend on exactly how the prompt was worded is real, and it is mostly concentrated in a minority of intents rather than spread evenly. But you cannot know which of your intents is the unstable one without asking it more than one way, and almost nothing on the market does that.
Why the churn half exists at all
An assistant re-plans retrieval on every request. It is not reading a fixed index of trusted sources and quoting from it, and the sources that survive every reading of a question look different from the ones that appear once. In the cells here the durable half is dominated by a small number of large destinations that show up across many unrelated questions, while the churn half is long, thin and mostly made of pages that fit one phrasing of one question on one day.
That is why the two halves need reporting separately rather than summing. Working on the durable half is a different job from working on the churn half, and only one of them is a job you can plan.
What this means if you are buying AI visibility measurement
The chart is not useless. It is reporting two different things added together, and the useful one is the smaller one.
- Ask what the number counts. A percentage of durable sources over total sources moves mostly with the total. Ask for the set.
- Ask how many runs, and whether that number is fixed. A comparison across weeks where the number of answered runs drifts is measuring the drift. We measured that at twenty five points.
- Ask whether the tool runs more than one wording per intent. 5 of 6 intents do not care, and finding the sixth is the whole value.
- And ask whether it separates the core from the churn, because six days moved half the set and left three quarters of the core alone. Those are two different stories about your visibility and only one of them is worth acting on.
If you want that separation on your own category rather than on ours, that is what our product does, and the method above is the one it uses.
The surface is not the bigger axis, and we predicted it would be
This study’s stated limit was one surface, so we closed it: the same 12 cells read on AI Mode the same day, so the pairing is inside the cell and only the surface changes.
The prediction, written down and committed before the first call, was that the surface would matter more than a week. It does not.
| Median overlap | |
|---|---|
| Same surface, six days apart | 50 % |
| Both surfaces, same day | 50 % |
That is the retirement threshold the design named, so the claim that mixing surfaces adds more noise than a month of drift is retracted here rather than quietly dropped.
What the Jaccard was hiding is better than what we predicted. A low overlap between a wide set and a narrow one can describe containment rather than disagreement, and this is containment: a median 94 per cent of the AI Overview set sits inside the AI Mode set, which is 1.52 times wider. 6 of the 12 cells are fully contained, and only 2 fall below half.
So the two Google surfaces are not answering with different sources. One is answering with a superset of the other, and the Jaccard punishes the narrow side for being narrow. If you are cited in AI Overview you are almost certainly cited in AI Mode, and the reverse is a coin toss.
The two cells that break it are worth naming because they are real disagreement rather than width: both phrasings of the flights question share only 3 of 7, with AI Mode answering from Skyscanner, Frommers and US News where AI Overview answered from Kayak, Expedia and Quora. Same day, same market, same question, two different travel sections.
Two Google surfaces nest. The third does not.
Containment is the finding above, and it came from two surfaces that share an index. That leaves a question a buyer needs answered: does one surface contain another because that is how retrieval works, or because those two are the same retrieval seen at two widths?
Gemini on the same 12 cells, same day, separates them. It does not nest.
| Same 12 cells, same day | Median containment |
|---|---|
| AI Overview inside AI Mode | 94 % |
| Gemini inside AI Mode | 42 % |
| Gemini inside AI Overview | 37 % |
That was the retirement threshold the design named, so the claim that a narrow set sits inside a wide one is now bounded to the pair it was measured on, and not published as a property of surfaces.
And the pair of numbers is the whole point. Gemini against AI Overview has a Jaccard of 0.25, and AI Overview against AI Mode has 0.50. Those look like the same kind of number, twice as much agreement in one case. They are not the same kind of number at all. One is a narrow set living inside a wide one; the other is two sets that genuinely do not share sources. Only the containment figure separates them, and no tool on the market reports it.
Gemini is also slightly narrower than AI Overview, a median 0.87 of its size, so this is not a width story. It is a different retrieval wearing the same company’s name.
What that means if you are spending money on measurement: watching AI Overview tells you most of what you would learn from AI Mode, and close to nothing about Gemini. Five of the eight comparable cells share less than half of Gemini’s sources with AI Overview. If your tool reports one Google number, ask which surface it came from, because the other one is a different market.
And it is Germany that breaks it, not America that is strange
If one Google surface contains the other because they share an index, that has to show up in any market. So the same eight intents in German, in Germany, on both surfaces, in one call each so the two readings cannot drift apart.
It showed up weaker, and we published that as the nesting is American. That was wrong, and a third market is what showed it. Spain, same eight intents, same design, lands at 80 per cent.
| Median containment, AI Overview inside AI Mode | |
|---|---|
| United States | 94 % |
| Spain | 80 % |
| Germany, all comparable cells | 50 % |
| Germany, the 5 cells read at three runs | 67 % |
The frozen threshold said 80 per cent or more would move the outlier from the United States to Germany, and Spain landed on it. So the correct sentence is the narrower one: in two of three markets, watching AI Overview tells you most of AI Mode, and Germany is where that stops being true. With two markets a single low reading looks like a rule about the other one.
The two numbers for Germany are both honest and they differ for a reason worth saying out loud. Containment is a fraction of the AI Overview set, so when both surfaces answer fewer runs both sets get narrower and the fraction falls. The two German cells that answered twice are the two lowest in that market. A run count is not a neutral setting; which direction it bends a figure depends on which side of the fraction it narrows. Spain replicated it without being asked to: its two cells at the lowest bars also sit below its own median, and its figure at three runs or more is the same 80 per cent, so the bar is not what puts Spain above Germany.
And the control is what makes the arm mean anything. If these answers had used the same sources as the American ones, this would be one market measured twice. They do not: joined question by question, the median overlap with the American source sets is 0.11 for Germany and 0.10 for Spain, neither running above 0.29. Three different sets of sources, and inside them the two surfaces relate differently.
One cell in each market returned no AI Overview at all, and it is the same question in both: running shoes for beginners, in German and in Spanish, three runs each, nothing. That is a surface declining to appear rather than a zero, so it is excluded and named, and the fact that it is the same intent twice says the declining is a property of the question and not of the country.
What this costs a buyer: you cannot assume, and the assumption is cheap to test. In the United States and Spain, one of the two surfaces is a good proxy for the other. In Germany it tells you half, and the German AI Mode set is 1.67 times wider. If you sell in more than one country, whether the two surfaces are one measurement or two is a question about each country, not about Google.
The third surface is separate everywhere, and the market gap is smaller than yesterday
Two things published above make a question neither answers. Gemini sits 37 per cent inside AI Overview in the United States, which reads as a property of the provider. And the relation between two surfaces changes by market, from 94 per cent to 50. So is Gemini a separate retrieval everywhere, or is that American too?
The same sixteen cells, read again with all three surfaces in one call each so the three run counts move together.
| Median containment | Gemini inside AI Overview | AI Overview inside AI Mode |
|---|---|---|
| United States | 37 % | 94 % |
| Germany | 33 % | 58 % |
| Spain | 20 % | 67 % |
Gemini is a separate retrieval in all three markets, and in Spain it is more separate than in the United States. The frozen threshold said either European market above 70 per cent would narrow that sentence to the United States; neither came close.
And the control is the part that costs us something. The pair that shares an index should reproduce yesterday’s market gap, and it reproduces the direction and not the size: Spain is still above Germany, by nine points instead of thirty. The bar does not explain it either, because a lower bar on both sides pushes containment down and Germany went up. So the honest version of the sentence one section above is that Spain sits above Germany on both readings, and how far above is not stable from one day to the next. The ranking is what we would publish; the gap is not.
Gemini is also the narrowest set in Germany, at 57 per cent of the AI Overview set, and exactly the same width in Spain. Being separate is not the same as being small.
One cell in each market again returned no AI Overview at all, and it is the same question it was yesterday: running shoes for beginners. That is now four readings, two languages, two days, and not one AI Overview. A surface declining to appear for a specific intent is the most reproducible thing in this study, and asking the same question in English turned it into its own measurement: AI Overview answered every software question we asked in Spain and no product question.
What this does not show
- Three Google surfaces and three markets. Retention is AI Overview in the United States only; the surface halves add AI Mode and Gemini on the same cells, and the market half repeats the containment measure in Germany and in Spain. ChatGPT and Perplexity retrieve differently and by this route ChatGPT returns no sources at all, so its zero would be the instrument.
- Three markets and three languages. Everything before the market section is the United States in English; the containment measure is repeated in Germany in German and in Spain in Spanish, and the cross-market control for both is in the register.
- Two points and then a third. Six days is one interval, not a trend, and nothing here says the churn rate is constant.
- 12 comparable cells. A median over 12 is a direction and not a precise figure, and the two cells excluded for answering at a lower bar are named in the register rather than quietly dropped.
- Nothing here is causal. We measured what changed, not why.
Frequently asked questions
What is the difference between the core and the rest?
The core is the set of domains cited in every reading of a question rather than in some of them. The rest appeared at least once. The distinction matters because the two behave differently over time: in this study the core retained 73 per cent over six days and the full set overlapped 50 per cent.
Does asking more times make the number more stable?
Not in the way people expect. More runs raises the number of sources seen at least once, so a strict “cited every time” statistic gets harder to satisfy as you sample more. That is why comparisons across time have to hold the number of runs fixed, and why we measured a twenty five point difference between mixed and matched bars.
Is a source that disappeared from the chart actually gone?
Often not. In this study, half of the full cited set differed after six days while three quarters of the core stayed. A domain that drops out of a weekly view may simply have been in the churn half the whole time, which is why the useful chart shows the core and the border separately instead of one line.
How many times should a tool ask each question?
Enough to tell a stable citation from a coincidence, and the same number every time it compares. What matters more than the exact figure is that the number does not drift between the weeks being compared, because then the comparison measures the sampling and not the market.
Does the wording of the prompt change the answer?
For most intents, no. 5 of 6 intents here returned an identical set from a keyword form and a question form on the same day. The sixth returned different sets on both dates it was measured, which is why the honest approach is to ask each intent more than one way rather than to assume the wording is safe.
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.