We Pooled 120 of Our Own Measurements. A Stability Score Mostly Reports How Broad the Answer Was.
Across 120 published measurements in four markets, the stable core stays a narrow band while cited breadth varies ninefold. The percentage moves anyway.
Every tool in this category reports how consistent your citations are, and so do we. We pooled 120 of our own published measurements, across four markets and five separate studies, to see what that percentage actually tracks. The set of sources that holds across every run is a narrow band almost everywhere: a median of 7 domains, with the middle eighty per cent sitting between 4 and 12. What varies is how many sources the answer cited at all, and that runs from four to thirty-five.
Divide a number that barely moves by a number that varies ninefold and you get a percentage that looks like a measurement of the category and behaves like a measurement of the question.
Disclosure: EchoWi sells AI visibility measurement, so a study concluding that a widely sold metric is partly an artefact is a study with an interest, and this one cuts against our own published percentages before anyone else’s. Nothing new was measured here. Every figure is recomputed from rows already in our measurement register, each with its query, market, surface, date and run count, and with that anyone can repeat the arithmetic against us.
The short version
- The stable core is a narrow band. Median 7 domains, tenth percentile 4, ninetieth 12. That holds across four markets and five studies.
- Cited breadth is not narrow. 4 to 35 domains, median 15.
- So the share falls as the answer widens. With two runs: 70 per cent where twelve or fewer domains were cited, 47 per cent in the middle band, 41 per cent above twenty. With three runs: 60, 29 and 26.
- The correlation is negative and modest, at −0.45 between how many domains were cited and what share of them held. Modest matters, and the limits section says why.
- The control is the cleanest number here. 8 single-run rows from four different studies all score 100 per cent, because with one run there is nothing for a source to be absent from.
- A sixth set flips the sign, and it is named below rather than dropped.
What was pooled, and what was left out
Five studies, all of them already published with their rows visible: the phrasing work, the travel booking questions, the sector agency sweep, the citation-slot categories and the answer-layer probes. Four markets, five surface combinations, 120 rows with two or more runs.
Two exclusions, and both of them cost the finding rather than help it.
Single-run rows are not in the correlation. With one run, every cited domain came back in every run, so the share is 100 per cent by construction. That is the artefact rather than a data point. They are counted separately instead, because a control that confirms the instrument behaves as expected is worth more than seven extra rows.
One collection is excluded because a reader cannot check it. Our register renders every measurement set except one, which is declared as an exemption in the data file because it backs no published figure. It happens to be the only one of the six whose sign goes the other way, at n=6. Removing it takes the exception out of the headline, which is exactly why this paragraph exists: with it included the correlation is −0.29 among two-run rows and −0.51 among three-run rows, and the conclusion does not change.
Run counts are kept apart rather than averaged. With two runs a source has to appear twice to be durable; with three, three times. Those are different bars, so mixing them would blend two definitions into one percentage, which is the mistake this whole article is about.
The numbers
How wide the answer was, against how much of it held:
| Runs | Domains cited | Rows | Median durable | Median share |
|---|---|---|---|---|
| 2 | 12 or fewer | 7 | 6 | 70% |
| 2 | 13 to 20 | 23 | 8 | 47% |
| 2 | more than 20 | 23 | 10 | 41% |
| 3 | 12 or fewer | 44 | 5 | 60% |
| 3 | 13 to 20 | 17 | 5 | 29% |
| 3 | more than 20 | 11 | 6 | 26% |
Read the “median durable” column down each block. With two runs it goes 6, 8, 10 while the answers behind it roughly triple in width. With three runs it goes 5, 5, 6.
The numerator is doing very little work. The denominator is doing almost all of it, and the denominator is a property of the question and the day, not of how settled the category is.
The spread, stated plainly:
| Across 120 rows | |
|---|---|
| Domains cited | 4 to 35, median 15 |
| Domains holding every run | median 7 |
| Middle 80% of the durable core | 4 to 12 |
| Correlation, cited against share | −0.45 |
| Correlation, cited against durable count | +0.48 |
Those two correlations are the finding in one line. More citation does bring slightly more stability in absolute terms, at +0.48, and it brings proportionally less, at −0.38. Both are true at once and only the second one reaches a dashboard.
Why this makes cross-category comparison unsafe
Put two categories side by side. One returns a tight answer from eight sources and six of them hold: 75 per cent. The other returns a sprawling answer from thirty sources and nine hold: 30 per cent.
On the percentage, the first category looks more than twice as settled. On the thing a buyer can act on, the second one has more durable sources to pursue, and its core is half again as large.
This is not hypothetical. It is the shape of the table above, and it is why our own studies have never ranked categories by stability share. The citation-slot work compares which domains hold a slot rather than what fraction of them did, and the cross-category study counts domains that span categories rather than averaging their quotas. That was a judgement call at the time. This is the measurement that justifies it.
The practical version: a stability percentage is comparable with itself over time on the same question, and it is not comparable between questions. Two categories, two markets, two wordings, or the same question on a day when the answer happened to cite more sources, are four different denominators wearing one label.
What this does not overturn
It would be easy to read this as “stability figures are meaningless”, and that is not what 120 rows support.
The durable set is real. It is the part of these measurements that has survived every check we have run against it: it survived re-reading the same question, it survived rewording, and where the quota has swung by fifty points the identity of the sources underneath moved far less. That is the whole reason our studies publish domain names.
A falling quota on one frozen question is still a signal. If the question does not change, the market does not change and the run count does not change, then a drop in share means either the core shrank or the answer widened, and both are worth knowing. What you cannot do is read one category’s number against another’s.
And the effect is a tendency, not a law. A correlation of −0.38 leaves most of the variation unexplained. Answer breadth is one input into that percentage, demonstrably, and it is not the only one.
What to ask a vendor, and it is the second question
We have published one question to ask any AI visibility vendor before anything else: do your repeated runs bypass the cache? If they do not, the trend line is flat by construction and you are buying a straight line.
This is the second one. Ask what the denominator of your stability score is, and whether it is held constant.
There are only a few possible answers. If the score is durable sources over sources cited, it moves when the answer gets wider, which is not a change in your visibility. If it is durable sources over a fixed universe the vendor defines, ask what is in that universe. If it is a raw count of durable sources, that is the most honest of the three and the least likely to be on the front of a dashboard, because it does not scale to a percentage.
A third question follows from the same place: how many runs is the score built from, and is that number shown next to it? Our two-run rows and our three-run rows produce different medians for the same bar, 8 against 5 in the middle band, purely because the bar moved. A percentage with no run count beside it cannot be read at all.
A companion study asks what happens to that percentage across two days rather than inside one: the same fifteen questions, asked again two days later, where the best and the worst row both moved when read a second time the same afternoon.
What this does not show
Observational, and pooled. These 120 rows were gathered for five different studies with different questions, verticals and surfaces. Nothing was randomised and nothing was held constant except run count, which is segmented. This describes a relationship in data we already had; it does not establish that widening an answer causes a lower share.
The bands are uneven. One cell has five rows in it. The direction is consistent across all six cells and across five of six studies, which is what carries it, not any single cell.
Our own studies are the sample. Every row is a question we chose to ask, in a market we sell in, on the surfaces available to us. A different sample of categories would produce different medians. What we would expect to survive is the shape: a narrow durable core against a variable denominator.
One surface family, mostly. Most rows come from Google’s AI Overview and AI Mode. We have published that source sets do not transfer between surfaces, so the medians here should not be read as universal across every assistant.
And no time dimension at all. Durable here means “came back in every run of one reading”. Whether a core that holds within a session also holds across weeks is the largest thing this register still cannot answer, and pooling more rows does not answer it. It needs the same frozen questions re-asked after time has passed, which is measurement rather than arithmetic.
Common Questions About Stability Scores
What is a stability score in AI visibility tools?
It is the share of sources cited for a question that come back every time the question is asked, and every product in this category reports some version of it. The number itself is well defined. The problem is that it is a ratio whose bottom half, how many sources the answer cited at all, varies from four to thirty-five across our measurements and is a property of the question and the day rather than of your visibility.
Why does a lower percentage not mean a less settled category?
Because the durable core barely moves. Across 120 published measurements its median is 7 domains and the middle eighty per cent sit between 4 and 12, whether the answer cited eight sources or thirty. So a category whose answers are broad scores a lower percentage while frequently having more durable sources in absolute terms, which is the opposite of what the number implies.
Can I compare my stability score against a competitor’s?
Only if you are both being measured on the same question, in the same market, on the same surface, with the same number of runs, on the same day. Change any one of those and the denominator changes underneath the percentage. Comparing the two lists of durable domains is safe and comparing the two percentages is not, which is why our own studies publish the domains.
How many runs should a stability figure be built from?
More than two, and the number has to be visible next to the figure. With two runs a single absence removes a source from the numerator entirely, so a small change produces a large swing. Our two-run and three-run rows give different medians for identically wide answers, 8 against 5 in the middle band, purely because the bar for “durable” moved. Tolerance is bought with runs, not with a cleverer statistic.
Is a 100 per cent stability score good news?
Check the run count first. Five rows in this register score 100 per cent and all five are single-run readings, from three different studies, where every cited domain is durable because there was nothing for it to be absent from. That is the instrument behaving as expected rather than a result. A genuine 100 per cent across three or more runs is worth having and worth re-measuring, because the best number in a study is the cheapest one to check and the one that changes most if it falls.
What should I track instead?
The set of domains, by name, and its size. A count of durable sources is honest, comparable with itself, and tells you where to work: those are the pages a buyer’s answer is being assembled from in your category. The percentage is the same information divided by something you do not control and cannot compare, and if you want one number, the count is the one that survives every check in this article.
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.