We Read One Question Eight Times. The Sources Did Not Move, the Answer Did, and That Breaks the Metric Everyone Uses.
Twelve domains cited in seven of eight runs and one in one. Cited-in-every-run is not measuring stability, it is measuring whether you caught the odd run.
Every stability figure on this site rests on the same bar: a source counts if it was cited in every answered run, three runs per question. Yesterday that bar flipped one cell from zero durable sources to twelve between two readings an hour apart, so we stopped adding cells and read a single question deeply instead. Eight answered runs. Twelve domains cited in seven of them, and one domain, google.com, cited in one. A second question gave the identical shape at five runs: twelve at four, google.com at one. The sources are not flickering. One run in every five to eight returns a different answer, and the strict bar drops all twelve at once when your sample contains it.
That makes cited-in-every-run a measurement of your sampling luck rather than of anything about the sources, and we have been publishing it.
Narrowed the same day. Four more questions on this frame and surface do not show the mechanism below, so it belongs to these two questions rather than to AI Mode: a durable count of zero turns out to be three different facts, and one category returns the same nine sources in every run. The arithmetic here is unchanged for the two questions it was measured on.
Disclosure: EchoWi sells AI visibility measurement, and this piece retires a statistic we use in nine published studies and that products in this category report as a headline number. The two questions, the surface, the market and the run counts are in our measurement register, so the arithmetic below can be checked and repeated against us.
The short version
- Twelve domains, one count. Read eight times, twelve sources were each cited in exactly seven runs. Independent flicker scatters counts; twelve landing on the identical number is one shared missing run.
- The odd run cites Google and nothing else.
google.comis the only domain present in the run the other twelve are absent from, in both questions. - So the bar is all-or-nothing. Sample three runs without the odd one and you get twelve durable sources. Sample three runs with it and you get zero. Nothing about the sources changed.
- The chance of catching it at three runs is 33 to 49 per cent on these two questions. That is not a margin of error, it is a coin flip deciding between an empty answer and a full one.
- What to report instead is the frequency, seven of eight, not a pass or fail against a bar that one run can veto.
What we actually did
Yesterday’s cross-surface study published that four cells returned no durable source on AI Mode. Re-asked an hour later, all four returned one, and United States payroll went from zero to twelve with the twelve being exactly the domains that had been in two runs of three the first time. We retracted the claim the same day, and the obvious next question was whether the bar or the sources were moving.
So instead of another cell, one question read as deeply as the tool would go. We asked for ten runs and got eight, which is stated because the denominator of every count below is the runs that happened rather than the runs requested.
Two questions, one shape
| Question | Answered runs | Domains at one short of every run | Odd-run domain |
|---|---|---|---|
| Best payroll service | 8 | 12 | google.com at 1 |
| Best accounting software | 5 | 12 | google.com at 1 |
The important column is the third one. Twelve sources each missing exactly one run, in both questions, with no source at six of eight and none at eight of eight.
If each source were independently unreliable, the counts would spread out: some at eight, some at six, some at five. They do not spread at all. Twelve domains sharing a single count is what one missing run looks like from the outside, and it is why the aggregate is worth reading rather than glancing at.
The control: the same two questions on AI Overview
If one run in five to eight answers differently, the obvious question is whether that is AI Mode or the measurement. So both questions were read the same way on AI Overview.
| Reading | Answered runs | Cited in every run | Cited once | Shared missing run |
|---|---|---|---|---|
| AI Mode, payroll | 8 | 0 | 1 | yes |
| AI Mode, accounting | 5 | 0 | 1 | yes |
| AI Overview, payroll | 9 | 6 of 6 | 0 | no |
| AI Overview, accounting | 10 | 4 of 20 | 12 | no |
Neither AI Overview reading has the shape. No domain sits alone at one run while everything else sits one short of the total, so the degenerate run is a property of AI Mode in these four readings rather than of the tool that measures both.
But AI Overview is not deterministic either, and the first reading nearly convinced us it was. Payroll returned six domains, all six cited in all nine answered runs, no tail whatsoever. That is as stable as a retrieval layer gets, and it was tempting to write that AI Overview simply does not move. The accounting reading killed it: four domains in all ten runs, one in nine, two in eight, one in two, and twelve domains cited exactly once each.
So the two surfaces break the same statistic in two different ways. On AI Mode a single run can empty the durable set. On AI Overview the durable set erodes as runs are added, because the bar asks for unanimity and every extra run is another chance to lose it.
What that means for our own published figures
This is the uncomfortable part. Our register has a United States accounting row on AI Overview with thirteen durable domains at three runs. Read ten times, the same question on the same surface has four.
Both are correct and they answer different questions. Thirteen is how many sources came back in all three runs of that reading. Four is how many came back in all ten of this one. The bar is a function of the run count, and it always was.
That does not invalidate the comparisons built on it, because every cell in the matched-category study was measured at three runs and is being compared against other cells at three runs. What it does retire is a reading of those numbers we never wrote down but which anyone would make: a durable-set count is not an estimate of how many sources are stable. It is a count of what survived a specific number of runs, and it belongs next to that number.
The arithmetic that turns this into a problem
If one run in n returns the odd answer, the chance that a sample of k runs contains it is 1 - (1 - 1/n)^k.
| Runs sampled | Payroll, odd run 1 in 8 | Accounting, odd run 1 in 5 |
|---|---|---|
| 3 | 33 per cent | 49 per cent |
| 5 | 49 per cent | 67 per cent |
| 10 | 74 per cent | 89 per cent |
Read that table the right way round, because it is counterintuitive. More runs make the strict bar more likely to return zero, not less. The bar asks for unanimity, and unanimity gets harder as the sample grows. A product that runs your prompts daily and reports how many sources were cited every single time is running a test that its own diligence guarantees will fail.
That is the opposite of how a stability metric is supposed to behave, and it is what our own three-run figures have been sitting on.
Why the zero and the twelve were both true
United States payroll on AI Mode, first reading: three runs, one of them the odd run, zero sources cited in all three. Second reading an hour later: three runs, none of them the odd run, twelve sources cited in all three.
Both readings are correct. Neither describes the sources. At 1 in 8, a three-run reading lands on zero about a third of the time, and the first one did.
This also explains a case we published four days earlier and could not explain then. In our cross-day work an email marketing question returned nothing durable one afternoon and seven domains an hour later, and those seven were exactly the sources that had been in two runs of three. We wrote at the time that a strict bar and a shifting retrieval produce an empty set without anything being broken. That was right and vague. The mechanism is a single degenerate run, and it is visible the moment you read deeply enough.
What we are changing
We are not deleting the durable-set figures. They are still the honest answer to “which sources came back in every run of this reading”, and every one of them is published with its run count next to it. What they are not is a property of the question, and our studies that treat a change in that number as a change in the world need the caveat this piece provides.
The number we will lead with is the frequency. Seven of eight is a fact about the source. Zero of three, when the three included the odd run, is a fact about the sample.
And the second question to ask a vendor now has a companion. After “do your runs skip the cache”, ask what their stability number does when the number of runs goes up. If it goes down, they are reporting unanimity, and unanimity is not stability.
What this does not show
Two questions, one surface, one market, one day. This is AI Mode in the United States in English on 12 August 2026. The two AI Overview readings above are a control on the same day and the same market, not a general claim about that surface, and none of this says how often any of it happens over weeks.
The shared run is an inference from counts, not an observation of runs. The tool returns how many runs cited each domain, not which ones. What makes the shared-run reading more than a guess is that all twelve counts are identical in both questions, and that a separate three-run reading returned all twelve at three of three, which cannot happen if their missing runs were spread out.
Both readings stopped early. We asked for ten runs and got eight and five. That is why the two odd-run rates, 1 in 8 and 1 in 5, are estimates from a single reading each rather than measurements, and why the probability table is a range and not a constant.
And one degenerate run is not necessarily one kind of run. We can see that google.com is the only domain cited in it. We cannot see whether it was a shorter answer, a refusal, or a different retrieval path, and this measurement was not designed to tell those apart.
Common Questions About Measuring Citation Stability
What does cited in every run actually measure?
On this evidence, whether your sample included the run that answers differently. Reading one question eight times, twelve sources were each cited in seven runs and one domain in the remaining one. A three-run sample that contains that run reports zero stable sources; one that does not reports twelve, and the sources were identical both times.
Does running more prompts make a stability score more reliable?
Not this kind of score. A bar that requires a source in every run gets harder to clear as runs are added, so more diligence produces a lower number rather than a firmer one. On our payroll question the chance of the strict bar returning zero rises from 33 per cent at three runs to 74 per cent at ten.
What should be reported instead?
The citation frequency for each source, with the number of runs beside it. Seven of eight is checkable and it is about the source. A pass or fail against unanimity can be flipped by a single run and is about the sample.
Why did one run cite only Google?
We do not know, and this measurement cannot say. What is visible is that in both questions there is exactly one run where the twelve usual sources are absent and google.com is present. Whether that run was a shorter answer or a different retrieval path is a separate question that needs per-run inspection rather than aggregates.
Does this mean your earlier studies are wrong?
The counts are right and the interpretation needed narrowing, which is why this is published rather than quietly fixed. A durable set is an accurate description of one reading. Treating a change in it as a change in what the engine cites is the part that does not survive, and yesterday’s cross-surface piece already carries the retraction that started this.
How would I repeat this against you?
Take one buying question, ask it eight to ten times on one surface in one market, and write down how many runs cited each domain instead of whether each domain cleared every run. If the counts cluster on a single number one short of the total, you have the same shape, and the bar everyone quotes is measuring your sample.
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.