One Category Cited the Same 13 Sources Every Time. Another Cited 49 and Repeated Almost None.
42 AI Overview measurements across six categories: stability ran from every source repeating to none. Re-measured a day later, one category flipped.
Ask Google’s AI Overview for the best AI coding assistant three times and it cites thirteen sources, and the same thirteen every time. Ask it which firm to hire for fractional CFO work three times and it cites eighteen, of which twelve appear exactly once.
Same surface. Same market. Same day. Same number of runs. The difference is entirely the category being asked about.
That matters because every AI visibility product on the market, ours included, sells you a number of runs. If the right number is three in one category and thirty in another, then a single sampling rate applied to every client is wrong for most of them, and nobody in this market publishes the figure that would tell you which case you are in.
How this was measured: Google AI Overview, via a live scraper rather than a model API, with country and language set explicitly on every call. Cache bypassed, every run a fresh upstream request, executed serially. Ten measurements, 42 runs in total, on 5 to 7 August 2026. We count cited domains, not brand mentions: they are different outcomes and mixing them is the most common error in this category. Every prompt, market, date and run count is recorded verbatim in our measurement register.
The result
| Category | Market | Runs | Domains cited | In every run | Exactly once |
|---|---|---|---|---|---|
| AI coding assistants | US | 3 | 13 | 13 | 0 |
| Flight booking | US | 3 | 6 | 6 | 0 |
| Noise-cancelling headphones | US | 2 | 4 | 4 | 0 |
| CRM for small business | US | 4 | 11 | 3 | 4 |
| AI visibility tools | Spain | 3 | 11 | 5 | 4 |
| AI visibility tools | France | 3 | 13 | 3 | 7 |
| AI visibility tools | Spain | 3 | 17 | 3 | 11 |
| Fractional CFO services | US | 3 | 18 | 3 | 12 |
| AI visibility tools | US | 4 | 23 | 3 | 14 |
| GEO agencies | Spain | 14 | 49 | 0 | 39 |
| value | |
|---|---|
| AI coding assistants, US | 100% |
| Flight booking, US | 100% |
| Noise-cancelling headphones, US | 100% |
| AI visibility tools, Spain | 45% |
| CRM for small business, US | 27% |
| AI visibility tools, France | 23% |
| AI visibility tools, Spain | 18% |
| Fractional CFO services, US | 17% |
| AI visibility tools, US | 13% |
| GEO agencies, Spain | 0% |
Read the last two columns together. At the top, every source the answer used, it used every time. At the bottom, forty-nine domains were cited across fourteen runs and not one appeared in all of them.
The comparison with no excuses in it
Different run counts make the middle of that table hard to compare, so here is the pair that has nothing to explain away.
AI coding assistants and fractional CFO services. Both United States, both English, both Google AI Overview, both three runs, both measured on 7 August 2026.
| AI coding assistants | Fractional CFO services | |
|---|---|---|
| Domains cited | 13 | 18 |
| Cited in all three runs | 13 | 3 |
| Cited exactly once | 0 | 12 |
One of these is fully reproducible at three runs. The other, at the same cost and on the same afternoon, tells you almost nothing at three runs: two thirds of what you would have seen was a one-off.
If you sold both of those companies the same monthly report, one of them would be getting a measurement and the other would be getting noise with a chart on it.
What actually predicts it, and what does not
The obvious guess is category age. It is wrong in both directions.
AI coding assistants is roughly three years old and perfectly stable. CRM software is thirty years old and returned three stable domains out of eleven. Newness does not cause instability.
The second guess is category size or money. Also wrong: flight booking is one of the most commercially contested queries on the internet and it returned six domains, all six stable.
What the unstable rows have in common is subtler, and we are naming it as a hypothesis rather than a finding because ten measurements cannot establish it:
Instability tracks the number of roughly interchangeable pages competing to answer. Fractional CFO firms, GEO agencies and AI visibility vendors have all produced a large body of near-identical commercial pages, each making a similar claim with similar authority. There is no obvious source to prefer, so retrieval assembles a different plausible subset each time. Where a category has a small number of clearly primary sources, whether that is Gartner and GitHub for coding tools or a handful of metasearch engines for flights, retrieval finds the same ones and stops.
If that is right, the counter-intuitive consequence is that more content in a category makes citations less stable, not more. That is worth testing properly and we intend to.
The two things this changes about measurement
One: your sample size is a property of your category, not of your budget.
The standard product in this market is a fixed number of runs per prompt per month, sold identically to everyone. Our own catalogue of verified pricing shows the range: Nightwatch gives thirty answers per prompt per month on every tier, Rank Prompt sells a credit pool you allocate yourself, and several vendors do not publish the figure at all.
None of those is wrong. What is wrong is choosing between them without knowing which half of this table you are in. Thirty answers a month is generous in a category like coding tools and thin in a category like professional services.
Two: a single report cannot tell absence from noise.
In the unstable categories, most cited domains appeared once. If your brand is one of those, a monthly report saying “cited” and next month’s saying “not cited” describes no change in the world at all. You would need many more observations before the difference between those two months meant anything, and the arithmetic on how many is unforgiving at small sample sizes.
This is the same shape we found when we re-ran a three-market study two days later: the aggregate reproduced almost exactly while the individual names churned almost completely. Stable totals, unstable membership.
The test you can run yourself
You do not need our data. You need yours, and this is the whole procedure:
- Take the prompt you actually care about.
- Run it three times, with country and language set explicitly, cache bypassed.
- List the domains cited in each run.
- Count how many appear in all three.
If most of them do, three runs is a real measurement in your category and you can spend the rest of your budget on breadth: more prompts, more markets, more surfaces.
If few of them do, three runs is a draw and not a rate. Spend on repetition instead, and treat any single monthly number, from any vendor, as provisional until you have enough runs to put an interval around it.
Do this once per category, not once per prompt. It is a property of the question space you sell into and it will not change week to week.
The finding we were not looking for
The CRM measurement was requested three times and Google returned an AI Overview only once. No error, no failure: two of the three requests simply produced a page with no AI answer on it. We re-ran it at four requests and got four.
That is a third state, and most reporting in this category collapses it into the second:
- You appear in the answer.
- There is an answer and you are not in it.
- There is no answer.
A tool that reports “0 mentions” for the month has told you nothing about which of the last two you are in, and they call for opposite responses. If a question often produces no AI Overview at all, your work is not to write a better page for it. It is to pick a different question, or to accept that the surface is not where that buyer is being served.
We have written before about answers that arrive with no sources attached, which is a related but distinct case: there the answer exists and cites nothing. Here the answer does not exist. Both are invisible in a visibility percentage.
We have since counted which domains recur across these categories rather than within them, and one platform reaches almost all of them: YouTube was cited in eleven of twelve measurements, in every run of eight.
We re-measured one row the next day and it inverted
The table above is ten sessions on named days. On 8 August 2026, one day after it was measured, we ran the CRM row again: the same prompt string, the same surface, the same market and the same number of runs.
| CRM software for small business, US, Google AI Overview | 7 August | 8 August |
|---|---|---|
| Runs | 4 | 4 |
| Domains cited | 11 | 7 |
| Cited in every run | 3 | 7 |
| Cited exactly once | 4 | 0 |
The category that supplied this article’s example of instability came back perfectly stable. Twenty-seven per cent became one hundred, overnight, with nothing changed at our end and nothing left to explain it away: identical prompt, identical run count, one day apart.
One note on how nearly this went wrong, because it is the same trap this article is about. The first re-measurement used “what is the best CRM for a small business?”, dropping one word from the registered prompt, and returned seven of seven at five runs. That looked like the same result and was not the same test. The comparison above uses the exact string in the register, because a prompt that differs by a word is a different prompt and we have published that rule elsewhere.
Three things follow, and the first two cost us something.
The stability ranking in the table above is a snapshot of sessions, not a property of categories. We wrote that this article was a photograph and not a series. This is what that sentence actually means when you cash it: a category can move from the unstable half to the top of the table in a day. Anyone using our table to decide how often to measure their own category should treat the row as one observation, which is what every row here has always been.
The mechanism we hypothesised is not supported by this. We suggested instability tracks the number of roughly interchangeable pages competing to answer. The population of CRM pages on the internet did not change materially between a Friday and a Saturday. Whatever moved was on the retrieval side, not in the corpus, so a stable count of competing pages cannot be the whole story.
And the practical advice survives intact, for a reason worth stating. Nothing here suggests measuring once is safe. It suggests the opposite: if a single session can report 27% and the next reports 100%, then a single session is exactly what you cannot rely on. The fix is unchanged and now better evidenced: repeat, on uncached runs, and publish the run count beside every figure.
We also added a category the original set did not cover, measured the same day:
| Category | Market | Runs | Domains cited | In every run | Exactly once |
|---|---|---|---|---|---|
| Project management software | US | 3 | 18 | 4 | 8 |
Eighteen domains at three runs with four surviving all three, which places it with the unstable half rather than with flights and coding assistants.
What this does not show
- One surface, so the engine was never a variable here. Google AI Overview only. Every number in this article is within-surface, which means nothing in it can attribute stability to the category rather than the engine: we never varied the engine. We have since asked four surfaces the same question and found that not one cited source appeared on all four, including the category at the top of this table, whose within-surface stability was perfect. Separately, two uncached ChatGPT runs of one buying question shared no cited domain at all, which is worse than anything in the table above, so treating these figures as a general property of AI answers rather than of one surface would be a mistake.
- Ten measurements. Enough to establish that the spread between categories is large and real. Not enough to rank categories precisely, and certainly not enough to prove the mechanism we hypothesise above.
- The run counts differ. A domain must hit 3 of 3 to be stable at three runs and 14 of 14 at fourteen, so the metric gets harsher as runs increase. That is why the argument rests on the same-run-count pairs rather than on the full table. It does not explain a gap between 0% and 67% one-offs at identical run counts.
- Nothing causal. This is observation. We changed nothing and measured what came back.
- Nothing about quality. Being cited is not being good. It says what the retrieval layer found, not what is worth buying.
- A photograph, not a series. These are single sessions on named days, and our own re-measurements show sessions disagreeing with each other. Treat every row as one observation.
Why we are publishing this
It argues against the simplest version of our own product. A vendor selling AI visibility measurement has an obvious incentive to say that measurement is straightforward, that a monthly number means what it appears to mean, and that the same package suits everyone.
The data says otherwise in more than half the categories we tested, including our own. We would rather publish the number that complicates the sale than defend a simpler one that a customer can disprove with three runs and a cent.
If you want the arithmetic behind sample sizes, it is in our piece on share of voice. If you want to see what each vendor publishes about how it measures, which is less than you would hope, it is in our comparison of every tool we verified.
Common Questions About AI Answer Stability
How many times should I run a prompt to measure AI visibility?
It depends on the category, and you can find out for yourself. Run the prompt three times with country and language set explicitly, then count how many cited domains appear in all three. In our measurements that number ranged from all of them to none of them. Where most sources repeat, three runs is a real measurement. Where few do, you need considerably more before a monthly figure means anything.
Do AI answers change between identical requests?
In some categories, almost not at all. Asked for the best AI coding assistant three times, Google’s AI Overview cited thirteen domains and the same thirteen each time. In others, heavily: fourteen runs on Spanish GEO agencies cited forty-nine domains and none of them appeared in every run.
Which categories give the most stable AI citations?
In our sample, established consumer and product categories: flight booking, noise-cancelling headphones and AI coding assistants all returned every cited domain in every run. The least stable were professional services and emerging B2B software, where most cited domains appeared exactly once.
Why would a newer category be more stable than an older one?
Age does not appear to be the driver. AI coding assistants is a young category and was perfectly stable; CRM software is decades old and was not. Our hypothesis is that stability tracks how many roughly interchangeable pages compete to answer the question, so a category with a few clearly primary sources returns them consistently while a crowded one returns a different plausible subset each time. That is a hypothesis, not a result.
What does it mean if my brand appears one month and not the next?
In an unstable category, quite possibly nothing. Most cited domains in those categories appeared in exactly one run out of three or four, so a change between two single observations is well within the noise. Before treating it as a real movement, check how many runs sit behind each number.
Does Google always show an AI Overview?
No, and this matters for measurement. One of our queries returned an AI Overview on only one of three requests, with no error on the other two. A report showing zero mentions cannot distinguish “there was an answer and you were not in it” from “there was no answer”, and those call for different responses.
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.