A Durable Source Count Is Not What It Looks Like. Here Is What Actually Predicts an Answer Set You Can Name.
Six questions read deeply, then the same intent asked three ways. Only one wording returns a source set you could hand a buyer, and it is the shallowest.
Yesterday we published a mechanism: read one buying question eight times and twelve sources come back in seven runs each, with one run citing nothing but Google, so a strict bar drops all twelve at once when your sample contains that run. We wrote it as a property of the surface. Today we asked four more questions on the same sentence frame, the same surface and the same market, and the shape appears in none of the 4 questions we added. CRM returns nine domains and all nine clear every run, twice measured. Web hosting returns 41 and nothing clears. Ecommerce returns 64 and nothing clears. So “zero durable sources” is at least three different facts wearing one number, and we had generalised from the middle case.
The useful part is not the retraction. It is that the three cases need opposite responses, and the number cannot tell them apart.
Disclosure: EchoWi sells AI visibility measurement, and this narrows a claim we published yesterday and a statistic we use across nine studies. Every question, with the surface and market it was asked in, its run counts and every domain with its count, is in our measurement register, so this can be repeated against us. The piece has grown well past the one surface and one market it started on, and the limits section at the end says exactly how far.
The short version
- The sweep: six questions, one sentence frame, one surface, one market, one day. Only the category changes. Everything after it widens one of those.
- CRM returns nine domains and all nine are cited in every answered run, for that string. Read again the same day, the same nine domains, again all of them. Unanimity is achievable, and bullet 7 says what it costs to reach it.
- Three questions return nothing durable, and not for yesterday’s reason. Password managers, web hosting and ecommerce show ordinary dispersion: no shared count, no lone Google, just a long tail.
- Yesterday’s shape appears in two of six questions. It is real and it is not the surface.
- The control that matters: CRM and accounting both answered five runs. Same depth, same surface, same market, same frame. Nine of nine against zero of thirteen.
- Read deeply, only one fixed set survives. That shape at three runs is usually the bar being shallow, and a cell that looked fixed in the United Kingdom has nothing left at seven runs.
- The same intent in three wordings: 9 cited and 9 kept, then 80 and none, then 100 and none. Same market, same surface, same day. A fixed answer set belongs to the surface and the string together, and asking the same thing on the other surface is what shows which.
- What predicts a nameable set is not the run count, it is how wide the answer cites. Readings that cite nine domains or fewer keep 56 to 100 per cent of them; readings that cite thirteen or more keep at most 40.
- And the widest category is not wide on the other surface. Three of the same strings on AI Overview cite 6 or 7 domains each, all of them durable. Ecommerce is 104 on AI Mode and 6 here, so the puzzle the rest of this page is about belongs to AI Mode rather than to the categories.
What we did
Yesterday’s reading took two questions on AI Mode in the United States and asked for ten runs each. Today we took the same sentence frame, “What is the best X for a small business?”, and ran four more categories the same way. Nothing changed except the category.
Every reading stopped before ten. That is stated because the denominator of every count below is the runs that happened, not the runs requested, and those denominators differ from row to row.
Six questions, three shapes
| Category | Answered runs | Domains cited | Cited in every run | Highest count |
|---|---|---|---|---|
| CRM | 5 | 9 | 9 | 5, by all nine |
| Payroll | 8 | 13 | 0 | 7, by twelve |
| Accounting software | 5 | 13 | 0 | 4, by twelve |
| Password managers | 4 | 19 | 0 | 3, by two |
| Web hosting | 7 | 41 | 0 | 5, by three |
| Ecommerce platforms | 4 | 64 | 0 | 3, by four |
Read the last column rather than the fourth. Four of these six rows report zero durable sources, and the last column says they got there three different ways.
CRM is unanimous. Nine domains, every one of them in all five runs, and no tail at all. Nothing was cited once.
Payroll and accounting are yesterday’s shape. Twelve domains sitting on the identical count one short of every run, and one domain alone in the run they all miss. One run differs; the other seven or four agree completely.
Password managers, hosting and ecommerce are dispersion. The highest count is reached by two, three or four domains, and everything else trails away. Ecommerce cites 64 domains across four runs, of which 52 appear exactly once.
Those last three do not have a degenerate run to blame. They have a retrieval that picks a different set each time.
The control, and why it is the important row
The obvious objection to CRM is depth: it answered five runs and payroll answered eight, so maybe CRM only looks unanimous because it was read less deeply.
Accounting answered five runs too. Same depth, same surface, same market, same sentence frame, same day. Accounting returns zero durable sources of thirteen cited. CRM returns nine of nine.
So run count does not explain the difference, and neither does the bar. Whatever separates them is the category.
We re-read the best number, because it was the best number
Nine of nine with no tail is the kind of result that is usually an artefact. Our own rule is to re-measure the number that would most change the conclusion if it fell over, so CRM was asked again the same day.
Three answered runs the second time. The same nine domains, checked by name, and all nine cited in all three runs. www.youtube.com, www.reddit.com, www.salesforce.com, slack.com, www.pcmag.com, www.uschamber.com, www.zoho.com, www.hubspot.com and fayedigital.com.
It survived. That matters more than the first reading did, because it means there are categories where this surface answers from a fixed set, and a buyer in one of them can be told which nine pages hold the answer to that question as it is worded. Two sections down, rewording it takes the set away.
And the fixed set belongs to the market, not the category
The CRM result invites one obvious question, so we asked it the same day: is a fixed answer set a property of the category, or of the market it was measured in? The same question, translated, on the same surface in Spain, France and Germany.
| Market | Answered runs | Domains cited | Cited in every run |
|---|---|---|---|
| United States | 5 | 9 | 9 |
| Spain | 4 | 50 | 0 |
| France | 6 | 42 | 0 |
| Germany | 6 | 63 | 1 |
One category, one sentence frame, 4 markets. The United States cites nine sources and keeps all nine. The other three cite between 42 and 63 and keep at most one.
One control before that table is read as a story about markets. Every row in it is AI Mode, and our publisher-layer count was measured on AI Overview. Asked on AI Overview in France the same day, the same question cites 8 domains and keeps 5 of them at four answered runs. On AI Mode in the same market it cites 42 and keeps none. That pair is not run-matched, though, and the difference matters: four runs against six, and more runs is more chances to cite something new. Matched properly, over the 10 cells where both surfaces were read three times, AI Mode cites more in 7 of them and the median ratio is 1.39, with the range running from 0.92 to 6.50. So the direction survives and the size does not: the surface does change how many sources are cited at all, by about half again in the typical cell rather than fivefold, and a slot measured on one surface still says nothing about the other.
Spain returns 50 domains across four answered runs and nothing cited in all of them. Read again, 34 domains across five runs with exactly one cited every time. Neither reading looks anything like the nine of nine with no tail that the United States gave.
The two Spanish readings do not even agree with each other. The domain the first reading had highest, at three runs of four, sits at one of five in the second. The domain the second reading has in all five runs was in two of four in the first.
And the American set held overnight, and then held against depth. Asked again the next day with six runs, the same nine domains came back, every one of them cited in every run, and asked a fourth time with eight, the same nine again. Depth is the axis that most threatens a fixed set, because more runs are more chances to cite something new, and here it bought nothing at all. That is 4 readings across 2 days and 22 answered runs, returning the identical set by name each time with no tail at any point. Our cross-day work on other questions found a median source overlap of 0.63; this cell sits at one.
And this one held when a second category was asked, which is the first thing today that did. United States business banking cites 7 domains and keeps all 7 at eight runs, the cleanest narrow cell here. Asked in Spain, in Spanish, on the same surface, it cites 14 and keeps 2. Same category, same surface, different market, and the shape flips exactly as CRM’s did.
Five answered runs against eight, and both halves of that cut in favour of the finding rather than against it: fewer runs can only cite fewer domains, and fewer runs make unanimity easier to clear.
Three claims in this piece were narrowed within hours of being published today, each by the next cell we measured. This one was tested the same way and did not move, so it is the one to lean on: a fixed answer set is a property of a category in a market, and the market is doing more of the work than the category.
So the sentence above needs a market attached to it. There are categories in a market where this surface answers from a fixed set. The same category one border away is assembled fresh, and a buyer there cannot be told which pages hold the answer, because there is no such set to name.
And it belongs to the wording and the surface, not only to the market
One question was left after that, and it is the one that decides whether any of this is usable. The nine domains held across 4 readings, 22 runs and 2 days, so they are not a fluke of one session. But every one of those readings asked the same string. So we asked the same intent, in the same market, on the same surface, on the same day, in different words: which CRM should a small business choose?
Seven answered runs. 80 domains cited, none of them in every run, the highest count 5 of 7, and 59 domains cited exactly once. That is the widest reading in the register, and it came from the cell we had just called the one place a buyer can be handed a list.
The fixed nine did not disappear, which is the part worth reading twice. Eight
of them are in this reading, and they sit at 1 to 3 runs of 7: slack.com at 3,
www.hubspot.com at 2, www.youtube.com, www.reddit.com,
www.salesforce.com, www.pcmag.com, www.zoho.com and fayedigital.com at 1
each. Only www.uschamber.com is absent. The same sources are still there. What
changed is whether the answer is assembled from them every time or from a long
tail that happens to include them.
Both rewordings so far used the same frame, so we asked a third that does not. “Which CRM should a small business choose” is still a question with a modal in it, and the widening could have belonged to that construction rather than to rewording. So we asked an imperative with no question at all: recommend a CRM for a small business. Eight runs requested and eight answered, the deepest read in the register. It cites 100 domains and keeps none, the highest count 5 of 8, and 66 domains cited exactly once. Four of the fixed nine are gone and the best of the survivors sits at 3 of 8.
| The same intent, three ways | Runs | Cited | Kept in every run |
|---|---|---|---|
| What is the best CRM for a small business? | 5 | 9 | 9 |
| Which CRM should a small business choose? | 7 | 80 | 0 |
| Recommend a CRM for a small business. | 8 | 100 | 0 |
Two structurally different rewordings both widen, so this is not one construction misbehaving.
But it is one surface misbehaving, and that is the bound this result needs. The imperative was asked again the same day, same market, same string, on AI Overview instead of AI Mode. It cites 12 domains and keeps 3, against 100 and none. Eight of the fixed nine are in it, three of them in every answered run.
So the collapse belongs to AI Mode rather than to rewording as such. On AI Overview the same rephrasing that shattered the AI Mode set leaves a narrow one standing. The honest reading of the three tables above is that the surface sets how wide an answer can get and the wording picks a point inside that range.
And that reading is already too strong, because the second intent does not reproduce it. The AI visibility question was asked on both surfaces the same day with the same string, and both readings answered five runs, so there is no depth to argue about. AI Mode cites 29 against AI Overview’s 21, and both keep 2.
That ratio is 1.38, and this register already published a median of 1.39 across ten run-matched cells. The fresh pair lands on our own median, which makes the CRM pair at 100 against 12 the outlier rather than the rule.
So no single axis dominates. The surface, the string and the market all move a cell’s breadth, the category does not, and the CRM result is where all three lined up at once. A reader who takes one number from this piece should take the median ratio and not the extreme.
The run counts are five against eight, and the two halves of that cut opposite ways, so both are worth stating. Fewer runs can only cite fewer domains, so the breadth gap is conservative and would survive more runs. Fewer runs also make unanimity easier to clear, so the 3 is the generous direction and we are not leaning on it. And read down that table rather than across it: the narrow row is one phrasing of three, and it is the shallowest of the three. A fixed answer set is not a place you can find and then own. It is one string among the many a buyer might type.
So a fixed answer set is a property of the surface and the string together, then of the market, and not of the category at all.
Outside English the same rewording does something else, and that is the more
useful half. French CRM on AI Overview, read twice, keeps 5 of 9 cited and 5
of 8. Asked with the same imperative frame, recommandez un CRM pour une petite
entreprise, it cites 18 and keeps 4. The durable count barely moved. Its
membership did: only www.shine.fr and foxeet.fr are in both, 2 of the 7
distinct domains the two readings keep between them.
So the English result, where rewording takes the durable count to zero, is not the general form. In French the count survives the rephrasing and the names underneath it turn over, which is the same asymmetry this register has measured before on other axes and the more dangerous of the two: a count that holds still reads like a set that holds still.
What tracking one prompt actually sees
The three CRM wordings are the same intent, the same market, the same surface and the same day, so their sources can be pooled. Between them they cite 131 distinct domains, and here is what each one on its own would have shown you.
| Wording | Domains cited | Share of the 131 |
|---|---|---|
| What is the best CRM for a small business? | 9 | 7% |
| Which CRM should a small business choose? | 80 | 61% |
| Recommend a CRM for a small business. | 100 | 76% |
Five domains appear in all three: www.youtube.com, www.reddit.com,
www.salesforce.com, slack.com and www.hubspot.com. 78 of the 131, or
60%, appear in exactly one of the three.
Do the arithmetic yourself: pool the domain lists of the three readings in our measurement register, count the distinct entries, and divide each reading’s own count by that total.
That is the practical version of everything above. A tool that tracks one prompt per intent, which is how this category is sold and how our own panel works today, shows you somewhere between 7 and 76 per cent of the sources building the answers to that intent, and you cannot tell which end you are on without asking the question a second way. The part that a single phrasing can reach is not most of it, and the part that survives all three is five domains, two of which are platforms rather than publishers.
The second intent bounds that range, and sharpens the finding rather than softening it. The AI visibility question was asked three ways too, and its three readings all answered five runs, so that curve carries no depth difference anywhere in it. They pool to 44 distinct domains, and each reading sees 41 to 52 per cent of them, nowhere near the 7 per cent floor of the CRM set. So the wide range above belongs to that intent, and specifically to it containing one unusually narrow phrasing.
What holds across both intents is the other number. 32 of the 44 domains, 73 per cent, appear in exactly one of the three wordings, against 60 per cent in CRM. And the curve replicates: one wording reaches a median 48 per cent of what the three find, two reach 75, against 61 and 79 in CRM. How much of the layer a single phrasing reaches moves a lot between intents. How much is reachable through only one phrasing does not, and it is the majority in both.
So what does a second wording buy? Taking the three CRM readings in every possible order, so the answer does not depend on which phrasing you happened to start with: one wording reaches a median of 61 per cent of what the three find together, two reach a median of 79 per cent.
The median is not the interesting part. The floor is. One wording ranges from 7 to 76 per cent depending on which one you picked. Two range from 62 to 99. The second wording buys you very little on a good day and rescues you entirely on a bad one, which is the shape of an insurance premium rather than a coverage upgrade.
And the third reads 100 per cent by construction, not by measurement. The denominator is the union of the wordings we ran, so the last one always closes it. That number is arithmetic, not evidence, and anyone quoting a coverage curve from a sample of prompts owes you the same caveat. What we can say is what one wording misses relative to three; what we cannot say is how much three miss.
And in the category this piece is about, neither wording works
Everything above is measured in proxy categories, which is deliberate: they are crowded, commercial and nothing to do with us. But the reader of this piece is usually shopping for an AI visibility tool, so we ran the same design on that question, on AI Overview in the United States.
Both readings answered five runs, which makes this the only wording pair in the register with no depth difference at all between its sides.
| AI visibility tools, United States | Runs | Cited | Kept in every run |
|---|---|---|---|
| What is the best AI visibility tracking tool for a brand? | 5 | 21 | 2 |
| Recommend an AI visibility tracking tool for a brand. | 5 | 18 | 3 |
Neither wording produces a set you could hand a buyer, including the “what
is the best” frame that produced the fixed nine in CRM. The two cited sets share
7 domains of 32 distinct, an overlap of 0.22, and of the five durable slots
across both readings exactly one domain holds in both: www.youtube.com.
That is worth sitting with if you sell in this category, and it is not a comfortable finding for anyone in it. A buyer asking an assistant which tool to use does not get a stable shortlist, gets a different one depending on how they phrase the question, and the single source that survives the rephrasing is a video platform rather than any vendor, review site or comparison page.
And it replicated in another category on the other surface. United States
business banking cited 7 domains and kept all 7 at eight runs, which is the
cleanest narrow cell in the register. Asked as which bank account should a
small business open, it cites 25 and keeps 3. Six of the original seven are
still there, from 1 to 5 runs of 5, and www.nerdwallet.com is still cited
every time.
The run counts differ, five against eight, and that cuts in favour of the finding rather than against it. Fewer runs can only cite fewer domains, and this cited three and a half times as many. Fewer runs also make unanimity easier to clear, and the durable share still fell from every domain to three of twenty-five. We already knew the strict-bar version of this from four wordings of a hotel question, where no domain survived all four. This is the stronger form: not that the durable set shrinks under rewording, but that a cell can go from nine of nine with no tail to eighty cited and nothing durable, one rephrasing apart.
And a three-run fixed set is usually not one
The shape that made CRM look special, every cited domain clearing every run, is not rare at three runs. Our matched-category study has cells like that in several markets, so we took the strongest one that is not the United States and read it deeply on its own surface.
United Kingdom payroll cited 8 domains at three runs and all 8 cleared. Read seven times the same way, it cites 19 and 0 clear, with the highest count at 6 of 7.
So the three-run version of that shape is mostly the bar being shallow. It says a category has few obvious sources, not that it has a fixed set, and the two look identical until somebody reads deeper. Of every cell we have now read at depth, the American CRM one is the only fixed set that survived.
That is also a caution about our own published tables. Every durable count in the matched-category study is at three runs, which makes it a fair basis for comparing cells against each other and a poor basis for telling a buyer which pages hold an answer.
We read one question six times in a row, and the count moved every time
Everything above reads a question once, deeply, or reads it again on another day. Neither answers a plainer question a buyer would ask first: if I run this tomorrow and the day after, am I looking at the same number?
So we read one question six times back to back, on one surface, three runs a reading, and then did it again with a second question in another market and another language. That is 13 calls and 38 answered runs, all on 14 August 2026. The design was written down before the first call, including what would kill it, because the register already carried an open question here and we did not want to answer it by choosing an analysis afterwards.
That open question was whether measuring repeatedly shrinks what you see. Five readings of one prompt on 10 August gave 18, 14, 11, 13 and 9 cited domains, they were taken in order, and a declining sequence taken in order cannot be told apart from noise. If it were real it would matter to everyone, because it would mean the act of monitoring degrades the thing being monitored.
It is not real. Six readings of the Spanish question gave 7, 6, 7, 9, 8 and 10 domains, and the highest of the six is the last one. The American question was read six times too, of which five completed three runs and one stopped early at two; the five gave 7, 15, 7, 7 and 7. Nothing shrinks.
The session does not explain it either, and that took one extra cell to establish. At position six of the first arm we inserted a question never asked in that session, same market, same intent, deliberately not one of the wordings read earlier. If the decline had been about being late in a session, the fresh question would have come back small. It returned 8 domains, which is inside the range of the question that had already been read five times. So the 10 August sequence was noise that happened to fall in order, and it is now written down as noise rather than left open.
What did move
The durable count. Six readings of one question reported between 2 and 5 durable domains, for the same prompt, the same surface and the same three runs, read one straight after another.
Pooled over 18 passes, 3 of the 16 domains seen missed at most one pass. Two more sit at 12 of 18, and those two are the reason the individual readings disagree: one reading promoted both to durable, another demoted one of the three that almost never miss. A three-run reading is not a noisy estimate of the durable set. It is a draw that lands on a different set each time, and it carries no signal about which of its members are the reliable ones.
The one fixed set, and its tail
The American question is the cell this corpus has already singled out as the only fixed set that survived a deep read. It survives this too, and more convincingly than before. Four of the five three-run readings returned exactly those 7 domains and nothing else, and the seven are the same seven, name for name, that a longer single read of the same cell returned on the same day.
The fifth reading returned those seven plus 8 domains that never appeared again. So the fixed set is a description of the core, not of the set. Read that cell twice and you have a real chance of drawing the one reading in six that makes a fixed category look like a churning one.
The rule we could not keep
The tidy explanation is that variation lives between calls rather than inside them: a batch of runs shares whatever state a batch shares, so more runs in one call buy less than the same number spread over several. It fits this arm perfectly. The longer single read saw no tail at all, and the tail turned up only when we asked again.
It dies against our own earlier rows. In the same category in Spain, France and Germany, single calls returned 50, 42 and 63 domains with 39, 28 and 36 of them seen exactly once, all inside one call. Tails plainly do occur within a call. What is true is narrower and stays: in this American cell, all 6 calls returned the core and 1 of them brought anything else.
We are writing the dead version down because it is the one the next person to look at this data will propose, and it was us.
Then the cheap version, on five questions at once
The section above reads two questions many times. Nobody buying a tool will do that. They will read a question, look at the number, and read it again next week. So we ran that instead, on five categories the register already covers, AI Overview, three runs a reading, two comparable readings each.
Two consecutive readings named the same durable sources in 2 of the 5 questions and different ones in 3. The median overlap between a question’s two readings is 0.67.
The single worst cell is the one worth the space. Accounting software cited 9 domains and 0 of them cleared every run. Read again minutes later, the same question cited 8 and all 8 cleared. Read a third time, the same 8 again. The sources had not changed at all: the eight were already there in the first reading, sitting at two runs out of three, because one run of that first reading cited nothing but Google. One run in three decided whether this category had no durable sources or eight of them.
That is the mechanism this article opened with, seen from the other end. There it explained why a deep reading returns zero. Here it flips a category between two readings a buyer would take a week apart and read as a trend.
The one that would pass unnoticed
Web hosting reported 5 durable domains in both readings, and they are not the same five. Two of the five changed places while the count stayed still. Anything reporting a count rather than a list would have shown no change at all, which is worse than showing a wrong number, because there is nothing to notice.
That is why the register keeps both figures per cell and why the table prints domains rather than totals.
What the agreements are made of, which is the uncomfortable part
Both of the perfect agreements are cells where every cited domain cleared the bar in both readings, 2 of the 5. Our design said in advance that this shape is the bar being shallow rather than a fixed set, and this article has already published that a three-run fixed set is usually not one. So the honest reading is not that two questions in five are stable. It is that the two questions that agreed are the two where the statistic had nothing to get wrong.
What this does not settle
Five questions, one surface, one day, and categories picked because we already had rows for them. The two readings sit minutes apart, so this measures agreement within a session and says nothing about a week. And two readings stopped early at two runs; they are in the register with their real run count and out of every paired figure, because a bar asking for unanimity is weaker at two runs than at three.
And a zero has two shapes, which we can now count
The lede of this piece says a strict bar drops twelve sources at once when one
run cites nothing but Google. That was a mechanism seen in one reading. It
leaves a shape behind, and the shape can be counted without measuring anything
new: google.com cited exactly once, and every other domain landing at exactly
one run short of the bar.
Counting every reading in our register with three or more answered runs, and leaving out the re-readings below because they were picked for being zero and would lift the share by construction: 80 readings, of which 18 return zero durable sources. 8 of those 18 have the near-miss shape and 10 are ordinary dispersion, where the best domain sits far below the bar rather than one run under it. The shape is not just Google turning up, because 14 readings cite it and only 8 of those match.
Those two zeros need opposite responses. A near miss says the sources are probably there and the bar caught a bad run. Dispersion says there is nothing to name. The reported number is the same.
Whether the near miss recovers, tested with a control
We re-read two cells of each shape at three runs. The control matters more than the test here, because this register has already measured that re-reading alone moves the statistic a lot, so a recovery means nothing unless something that should not recover fails to.
| cell | shape | durable share on re-read |
|---|---|---|
| Password managers, United States | near miss | 47 per cent |
| CRM, Germany | near miss | 3 per cent |
| Ecommerce, United States | dispersed | 3 per cent |
| Web hosting, Spain | dispersed | 0 per cent |
The best case is clean: password managers cited 19 domains both times and went from none clearing every run to nine. Nothing about the sources changed; the bar did. On AI Overview the day before, an accounting cell did the same thing and went to every cited domain clearing.
We took each arm to 5 cells, and the arms still overlap at the bottom. The near-miss shares are 2, 3, 47, 100 and 100 per cent, a median of 47. The dispersed shares are 0, 0, 0, 3 and 6, a median of 0. The lowest near-miss cell still sits under the highest dispersed one, so a rule reading “near miss means the sources are there” is wrong twice in five.
What separated is the other direction. No dispersed cell cleared 6 per cent, and 3 of the 5 near-miss cells came back at 47 per cent or more. The information is one-directional: a dispersed zero predicted another near-zero every time, and a near-miss zero predicted nothing. That is weaker than the symmetric claim we would have preferred and it is what ten readings support.
One caveat belongs on the ceiling of that arm rather than under it. 2 of the 10 re-readings returned every domain they cited as durable, both near-miss cells at 100 per cent, and this corpus already calls that a shallow bar rather than a fixed set. Part of that recovery is a narrower answer and not sources coming back.
The practical form is cheap enough to be worth it anyway. If your zero has the near-miss shape, read it again before acting on it. One more reading costs almost nothing and is the difference between “no source holds this category” and “nine do”.
Three runs is a weaker bar than the four to nine the original readings used, so re-reading favours recovery in both arms. That is why the only thing quoted above is the contrast, and why the contrast is reported as failing to separate.
What actually predicts a nameable answer set
Reading every deep measurement in the register together, the thing that separates a cell with a fixed set from one without is not the run count. The medians by run count go 78, 0, 3, 2, 0, 0, 100 and 20 per cent, which is no pattern at all.
What separates them is how many domains the answer cites in the first place. Of 59 deep readings, 14 cite nine domains or fewer, 40 cite thirteen or more, and five sit between them: the narrow ones keep between 56 and 100 per cent of what they cite, the wide ones keep at most 40 per cent. Depth does not explain it either, because the deepest narrow reading is 9 runs and the deepest wide one is 10.
And the predictor holds still, which is the part that decides whether it is usable. A predictor that wanders between days is not one, so the widest cell of the sweep was read again a day later with the same prompt, market, surface and run count. It cited 64 domains one day and 63 the next, and kept none either time.
The sources underneath did not hold still at all: 24 domains of 103 distinct appear in both readings, an overlap of 0.23. So the number reproduces while almost the whole membership turns over.
That asymmetry runs through this register, and everywhere else in it the stable number was the misleading one: a durable share that held while its sources changed completely, a stability quota that moved fifty points on a numerator that never left seven. This is the one place it works the other way. Breadth is what is being claimed, and it is a property of the cell rather than of which pages happened to come back.
So we predicted before measuring, which is the only way this stops being a description. United States business banking cited four domains at three runs. The pattern says a cell that cites few keeps them, so it was read eight times: it cites 7 and keeps all 7, with no tail. That is a fourth cell on the narrow side, in a market, category and surface combination that was not there before.
Two things this is not. The narrow side is 10 distinct cells against 29 on the wide side, so the side carrying the claim is the thin one. And a three-run cited count does not predict the deep one: United Kingdom payroll cited eight at three runs and nineteen at seven, which is what put it on the wide side. You learn which kind of cell you have by reading it deeply, not by counting a shallow reading.
The obvious confound is the surface, and it does not survive the test. AI Overview cites fewer domains than AI Mode, so narrow could be one surface wearing another name. Split by surface, the gap holds inside each: on AI Overview the narrow readings keep 56 to 100 per cent and the wide ones at most 40, and on AI Mode the narrow keep 100 and the wide at most 25. That test also shows where this is thin, which is worth saying in the same breath: the narrow side of AI Mode is 1 cell and the wide side of AI Overview is 8.
Then the prediction was aimed at the case designed to break it. United Kingdom CRM is narrow at three runs, seven domains, and none of them durable. The breadth pattern says a narrow cell keeps what it cites; the three-run reading says it keeps nothing. Read eight times it cites 6 and keeps all 6. The shallow count was not merely shallow there, it was inverted: the same cell reads as nothing stable at three runs and as six sources every single time at eight.
And the inversion replicated in a category we had never read. United States VPN was the same setup, six domains cited at three runs with none durable. Read eight times it cites six and keeps all six. Two independent cells, both reporting nothing stable at three runs, both returning a complete fixed set when read deeply.
What this does to yesterday’s claim
Yesterday’s piece said the sources do not flicker and that one run in every five to eight answers differently. Both halves are still true of the two questions it measured, and its probability arithmetic is unchanged for them.
What it should not have implied, and what a reader would reasonably take from it, is that this describes AI Mode. It describes 2 of 6 questions we have now read deeply on AI Mode in the United States. The other four split two ways, and one of them has no odd run at all.
This is the second time in two days that a finding of ours has needed narrowing when more questions arrived, and the shape of the mistake is identical both times: a clean mechanism that explains the batch that suggested it, published before it was checked against a batch that did not.
What to report instead of a durable count
A durable count of zero is not a finding on its own. What separates the three cases is the distribution of counts, and it is one line of arithmetic:
- Every cited domain at the maximum, as in CRM, means a fixed answer set. Work on those pages.
- Almost every domain one short of the maximum, one domain alone below, as in payroll, means a single odd run. The durable set is real, and your sample missed it. Re-read before concluding.
- Counts trailing away with a large one-run tail, as in ecommerce, means the answer is assembled fresh each time. There is no stable set to work on, and reporting zero is correct but useless without the shape beside it.
Any product reporting “durable sources: 0” for the second and third of those is reporting the same number for a near miss and for an absence.
Where the difference is not: the ranked page
The obvious explanation for CRM answering from a fixed set and ecommerce not is that their search results differ, one settled and one crowded. Testing that with the instrument that produced the puzzle would be arguing in a circle, so it needs a second instrument that knows nothing about the first: Google’s own organic results for the identical string, read the same day.
Six categories, the same string put to both layers, United States in English.
| Category | AI Mode runs | Cited domains | Also in the organic top 20 | Cited from off the ranked page |
|---|---|---|---|---|
| CRM | 8 | 9 | 8 | 1 |
| Payroll | 8 | 13 | 10 | 3 |
| Accounting software | 5 | 13 | 6 | 7 |
| Password managers | 4 | 19 | 7 | 12 |
| Web hosting | 7 | 41 | 8 | 33 |
| Ecommerce platforms | 9 | 104 | 7 | 97 |
The count of cited domains that also rank sits between 6 and 10 in every one of the six. What runs from 1 to 97 is the column beside it.
The percentage version of that, a share falling from 89 per cent to 7, is not a second finding and should not be read as one. Once the numerator is near constant, the share falls by arithmetic. The observation is the near constant numerator.
There is also a comparison this table does not make. The organic page’s own distinct-domain count is 13 to 16 across all six, which looks like a flat result and is not one: a top-twenty page is bounded at twenty by construction, so that flatness belongs to the instrument. Only the overlap is being read here.
The control that lets the categories be compared
The rows are read to different depths, four answered runs to nine, and breadth grows with runs. If the overlap grew with runs too, the table would be measuring depth. The widest cell was read twice for exactly that reason: 64 cited domains at four answered runs, and 104 at nine. The overlap with the ranked page was 7 both times, and the same seven domains by name. All 40 of the new domains came from off the page.
The same gap read from the other end
Turned around, the table says something a vendor pays for directly. Between 6 and 9 of the ranked domains are never cited at all, in every category. In ecommerce that includes four of the platform vendors themselves: Shopify, Wix, BigCommerce and Sellfy all rank in the top twenty for that exact string, and none of the four is cited in nine runs.
Cited is not named. This register keeps those two outcomes apart everywhere, and the distinction carries weight here: an answer can recommend a product at length while citing somebody else’s page about it. What is measured is whose page was cited.
So a page-one position buys a seat at the answer table whose size barely moves by category, and what moves is how big the rest of the table is. In CRM that seat is 8 of 9 sources. In ecommerce it is 7 of 104. The same ranking is worth two very different things, and which of the two you are in can be measured before spending anything on it.
What this still does not answer is why. It says where the difference is not. It is not on the ranked page, because the ranked page contributes roughly the same amount everywhere. Whatever makes a category answer from a fixed set is happening in what the surface reaches for beyond it, and this measurement does not reach there either.
And one control from outside the buying questions. Every cell above asks for the best product in a category. Asked instead how to rank in AI Overviews, which is a different intent in a different vertical, AI Mode over 6 runs cited 36 domains and 8 of them were in the organic top 20. That is 53 per cent of the ranked page, the median of the twelve. Seven ranked domains went uncited in six runs, Ahrefs at position five among them.
And a control on the other surface, which narrows all of this. Every cell above is AI Mode. Three of the same strings put to AI Overview on the same day, in the same market, cite 6 or 7 domains each, every one of them in every answered run, and draw 21 to 33 per cent of their ranked page. Ecommerce is the sharpest: 104 cited domains on AI Mode and 6 on AI Overview, for the identical question. So the category breadth this whole section is about is a property of AI Mode, and on AI Overview the wide category is as fixed as the narrow one.
The same design in a second market
Everything above is one market, and this register’s own rule is that a result measured in one market is a result measured in one market. So the six categories were run again in Spain, in Spanish, against Spain’s own ranked page for the identical Spanish string, on 14 August 2026.
Spain answers wider everywhere. Its narrowest category cites 25 domains where the narrowest in the United States cites 9, so nothing in Spain looks like the fixed set of nine.
| Category | Cited domains | Also in the organic top 20 | Share of the ranked page |
|---|---|---|---|
| Payroll | 25 | 6 | 67 per cent |
| Password managers | 27 | 6 | 40 per cent |
| Accounting software | 39 | 8 | 50 per cent |
| Web hosting | 50 | 4 | 50 per cent |
| CRM | 77 | 12 | 75 per cent |
| Ecommerce platforms | 114 | 9 | 56 per cent |
Across all 18 cells in the 3 markets, the cited set runs from 9 to 114 and the count of cited domains that also rank runs from 4 to 12.
The last column is what the three markets should be compared on, and it is the reason it exists. A raw overlap count is capped by how many domains the ranked page holds at all, and two of the Spanish pages returned nine and eight organic domains rather than fifteen. A share of the page each cell actually had is not capped that way. In the United States that share runs from 44 to 63 per cent, in Spain from 40 to 75, and in Germany from 27 to 63. The first two markets both land just above half. Germany lands at 44, the first market to sit below it, which is the kind of thing a third market is added to find.
So about half of page one gets cited, in the three markets, whether the answer draws on nine sources or a hundred and fourteen.
Two things this run does better than the first. Every Spanish cell answered exactly six runs, so the comparison across categories is run-matched, where the United States rows ran from four to nine answered runs and needed the depth control to defend them. And the finding no longer rests on one market, which is what this register asks of a claim before it is stated generally.
What eighteen cells cannot settle is whether that share moves with how wide a cell cites. What they show is the two ranges: a cited set spanning 9 to 114, and a share of page one spanning 27 to 75 per cent.
And a third market, read shallower
Everything above is two markets read at six runs and more. The same six categories were run again in Germany, in German, on 14 August 2026, with the same design and the same string put to both layers. All six answered three runs of three.
| Category | Cited domains | Also in the organic top 20 | Share of the ranked page |
|---|---|---|---|
| Accounting software | 11 | 7 | 44 per cent |
| CRM | 12 | 5 | 29 per cent |
| Payroll | 13 | 10 | 63 per cent |
| Password managers | 16 | 4 | 27 per cent |
| Web hosting | 45 | 9 | 60 per cent |
| Ecommerce platforms | 45 | 4 | 44 per cent |
Read the middle column, not the left one. These are three-run readings where Spain ran six, and this piece has already shown that a three-run cited count does not predict the deep one: United Kingdom payroll cited eight at three runs and nineteen at seven. So a cited span of 11 to 45 is not comparable with Spain’s, and the narrowness on the left is at least partly depth.
The column that is comparable is the one this study’s own depth control defends. Reading the widest American cell at four runs and again at nine took its cited set from 64 to 104 and left the count that also ranks at 7 both times: cited moves with depth, the overlap does not. In Germany that overlap runs from 4 to 10, inside the 4 to 12 the first two markets had already set. A third language moved neither end of it.
What Germany does move is the share of the ranked page, downwards. Its floor is 27 per cent against 40 in Spain, and its median is 44, the first of the three markets to land below half.
And the other surface points the other way here, which the American control did not predict.
Three of these German strings went to AI Overview on the same day, matching the three categories already controlled in the United States. There AI Overview draws less of the ranked page than AI Mode, 21 to 33 per cent against 44 to 63. In Germany it draws more, in all three: 59 against 29, 50 against 44, and 89 against 44.
Three cells that clean are usually the sentence rather than the market, so the same three were asked again in German a different way, on both surfaces. AI Overview came back at 53, 59 and 47 against AI Mode’s 35, 53 and 47. Over the two wordings that is 6 pairs, 5 with AI Overview above, 1 exact tie and 0 the other way.
So the direction is not the wording. The size is: ecommerce went from a 45 point gap to none at all. What these pairs support is that in Germany AI Overview was never the surface leaning less on page one, and not any particular distance.
Two limits, because this is the kind of contrast that invites a bigger claim than it can carry. The American side is three cells at one wording, so the half of the comparison that has not been repeated is the half the contrast rests on. And the American strings are English where the German ones are German, so that half still moves country and language together, which is exactly the confound the German re-wording removed on the German side only.
What this does not show
Our AI Mode counts are floors, not counts. Across 5 questions in 3 markets, an AI Mode answer prints numbered citation markers in its body and the structured list carries anywhere from 21 per cent of them to all of them, while AI Overview’s carried all of them in every cell. We first measured this on two cells, read it as a constant half, and a third market returned an AI Mode answer carrying every marker it printed. So every AI Mode breadth figure here is at least what it says, by an amount that varies per cell and that we cannot pin down. That happens to be the safe direction for this article: where we report AI Mode citing far wider than AI Overview, an under-captured AI Mode makes the real gap larger rather than smaller. What it rules out is correcting any of these numbers upward by a fixed factor, which the two-cell version invited.
It outgrew its own first limit, and this is the current one. This began as AI Mode in the United States in English on one day, and that sentence stood here after it had stopped being true. The readings now span 2 surfaces, 5 markets, 4 languages and 4 days. Every section says which it used, and no figure here spans more than the section it sits in.
The categories are not a sample of anything. They were chosen because our earlier work already covers them, so they are the categories we can compare against, not a random draw from the economy.
Answered runs vary from four to eight and we do not control that. The tool stops early and reports how many answered. A bar that asks for unanimity is harder at eight runs than at four, so rows are not interchangeable and the comparison that carries weight is the one where the depths match.
Why CRM is fixed and ecommerce is not, we still cannot say. The section above rules out one answer and only one: it is not that the ranked page differs, because the ranked page contributes six to ten sources whatever the category. That leaves the mechanism where the measurement cannot reach, in whatever the surface retrieves beyond page one, and a plausible story about small pools and crowded ones remains a hypothesis for a bigger study rather than a finding.
Nothing here is causal and nothing forecasts. These are readings, not an intervention with a control.
Common Questions About Durable Citation Sets
Does zero durable sources mean the engine is unstable?
Not by itself. In our six readings zero arrived two different ways: one where twelve of thirteen sources agreed and a single run differed, and one where dozens of sources each appeared once. The first is a near miss and the second is an absence, and the durable count is identical for both.
Can an AI answer return the same sources every time?
Yes, in some categories. Asking for the best CRM for a small business on AI Mode returned nine domains with every one cited in every answered run, and a second reading the same day returned the same nine by name. That is a cell where a buyer can be shown exactly which pages hold the answer to that question as it is worded. The same intent asked in different words, the same day on the same surface in the same market, cited 80 domains with none of them durable, so the list belongs to the phrasing before it belongs to the category.
Does ranking on page one get you cited in an AI answer?
Partly, and by an amount that barely changes with the category. Putting the identical string to Google’s organic results and to AI Mode on the same day across six categories, the number of cited domains that also rank in the organic top 20 landed between 6 and 10 every time, while the cited set itself ran from 9 to 104. So a page-one position contributes a roughly fixed number of sources, and what decides whether that is most of the answer or a fraction of it is how much the surface cites from off the ranked page. In the narrowest category it was 8 of 9 sources; in the widest, 7 of 104.
Does this hold on AI Overview?
No, and that is the most important limit on this page. Three of the same strings put to AI Overview on the same day cite 6 or 7 domains each, every one of them in every answered run, and draw 21 to 33 per cent of their ranked page. The clearest case is ecommerce: 104 cited domains on AI Mode and 6 on AI Overview for the identical question. The wide-versus-narrow split this page is built on is a property of AI Mode, so a figure measured on one surface says nothing about the other.
Why did your earlier article say one run in eight answers differently?
Because that is what its two questions did, and it remains true of them. Four more questions on the same frame and surface do not show it, so the mechanism belongs to those questions rather than to the surface. This piece is that correction, published rather than quietly applied.
How many runs should a stability figure use?
More than three, and the number has to be printed next to the figure. A bar that asks for a source in every run gets harder as runs are added, so the same question honestly measured at three runs and at ten produces different counts. In our register one question gives thirteen durable domains at three runs and four at ten.
What should I ask a vendor about this?
Three things now. Whether their runs skip the cache, how many wordings they track per intent, and what their stability number does when the run count goes up. Add a fourth if they report a zero: ask whether it came from one odd run or from a hundred sources appearing once, because they will not be able to answer without the count distribution.
How would I repeat this?
Take one buying question, ask it four to ten times on one surface in one market, and write down how many runs cited each domain instead of whether each domain cleared every run. Then do it for a second category in the same session. If the two distributions have different shapes, you have reproduced this, and the durable count alone would have hidden it. Then ask the first question again in different words, on the same surface on the same day, and pool the two domain lists: that is the step this piece turns on, and it is the one the method above does not reach on its own.
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.