We Asked AI for the Best Agency Nineteen Times Across Four Markets. Most of What It Cited Was Gone Next Run.
Nineteen buying questions, four markets, three AI surfaces, 46 runs. Of 408 cited sources, 58% appeared exactly once and about a third held across every run.
We put nineteen “best agency for X” questions to three AI surfaces across four markets. They cited 408 sources between them. Two hundred and thirty eight of those, 58%, appeared in exactly one run and were gone the next.
That is the number an agency needs before it puts an AI visibility figure in a client report, because it decides what the figure means. A source cited once is not a position in an answer. It is a coin landing heads.
Disclosure: EchoWi sells AI visibility measurement and competes in this category. That is a conflict, and the answer to a conflict is a method you can check rather than a promise to be neutral. Everything needed to repeat this, and to repeat it against us, is below: the questions, the surfaces, the markets, the run counts and the dates.
The short version
- Of 408 cited sources, 238 appeared exactly once. That is 58%. It was also 58% when this study was four questions smaller, in two fewer markets, which is the most reassuring thing in it.
- About a third held across every run. 129 of 408, 32%. That third is the only part of an AI answer worth building a strategy around.
- The number moves 10 points on run count alone. Questions that ran twice showed 37% stable; questions that ran three times showed 27%. Same method, same surfaces, same week.
- The market changes the answer, but not reliably. Of four comparisons matched on both vertical and run count, two differ by 15 and 35 points and two differ by 3 and 2.
- You do not get the runs you ask for. Ten of the nineteen questions rest on two runs or one, because the measurement stopped early and said so. Recording what ran instead of what was requested is the difference between a denominator and a wish.
What we measured
| Questions | 19, each of the form “what is the best marketing agency for [vertical]”, in the market’s own language. 18 carry a stability figure; one returned a single run and is excluded from every percentage |
|---|---|
| Verticals | Hotels, law firms, clinics, ecommerce, healthcare, B2B SaaS, white label SEO, higher education |
| Surfaces | Google AI Overview, Google AI Mode, Gemini, pooled per question |
| Markets | United States and English, Spain and Spanish, France and French, Germany and German, each set explicitly |
| Runs | 46 in total, 45 in the analysis, 1 to 3 per question, recorded per question rather than assumed |
| Cache | Bypassed. Every run is a fresh upstream call |
| Dates | 9 and 10 August 2026 |
Three things about that table decide whether any of this means anything.
The market is set explicitly on every run. Measurement defaults to the United States and English. A French question left on the default gets answered by the American market and nothing in the output says so, which would have produced four suspiciously similar results and a completely wrong conclusion.
The run count is what actually ran. Four questions stopped at two of three requested runs and the tool said so. Writing three there would have invented a third of a denominator, and inventing denominators is the failure this whole register exists to prevent. Every question below is a row in the measurement register with the count that came back, not the one we asked for. It also means the two-run and three-run questions are not directly comparable, which turns out to be the most useful thing in the study.
ChatGPT is not here. It is the surface people ask about most and the one we can least defend measuring this way, because the route we use returns no source list for it. A zero from an instrument is indistinguishable from a real zero, so it is left out rather than reported as absence.
How much of an AI answer actually holds still
Pool the three surfaces per question and you get a set of cited domains. Split that set into the ones present in every run, the ones present in exactly one, and the ones in between:
| Sources | Share of 408 | |
|---|---|---|
| Cited in every run | 129 | 32% |
| Cited more than once, not every time | 41 | 10% |
| Cited exactly once | 238 | 58% |
| cited in every run | distinct sources cited | |
|---|---|---|
| US · law firm, 2 runs | 13 | 20 |
| ES · hotels, 2 runs | 9 | 16 |
| FR · clinics, 2 runs | 8 | 16 |
| US · B2B SaaS, 2 runs | 7 | 19 |
| US · hotel, 3 runs | 6 | 17 |
| ES · ecommerce, 3 runs | 9 | 28 |
| US · white label SEO, 2 runs | 8 | 25 |
| FR · law firms, 3 runs | 7 | 22 |
| FR · ecommerce, 2 runs | 9 | 29 |
| US · healthcare, 2 runs | 6 | 20 |
| US · ecommerce, 3 runs | 7 | 24 |
| DE · law firms, 2 runs | 7 | 25 |
| FR · hotels, 3 runs | 7 | 26 |
| US · higher education, 3 runs | 6 | 23 |
| ES · clinics, 3 runs | 5 | 22 |
| DE · hotels, 2 runs | 5 | 24 |
| DE · clinics, 3 runs | 6 | 29 |
| ES · law firms, 3 runs | 4 | 23 |
Not one of the eighteen questions had a stable majority. The best was a United States law firm question at 13 of 20, and that one ran twice. It is also the one that fell to 7 of 24 when we asked it again the next day, which is a section of its own further down. The worst was a Spanish law firm question at 4 of 23, and that one ran three times. Which brings us to the part that changes how you read every number in this category.
The run count moves the number, and here is by how much
“Cited in every run” is an easier bar to clear over two runs than over three. That is arithmetic, not a finding. What is a finding is that we can put a size on it from inside one study, because the two groups were measured the same week with the same method:
| Questions | Stable share | Cited exactly once | |
|---|---|---|---|
| Ran twice | 9 | 37% | 63% |
| Ran three times | 9 | 27% | 54% |
| value | |
|---|---|
| Questions run twice | 37% |
| Questions run three times | 27% |
Ten points. On a slide, 37% and 27% are two different stories about a client’s position, and the difference between them here is not the client, the market, the vertical or the week. It is how many times somebody pressed the button.
This is the practical consequence, and it is worth being blunt about it. A visibility figure quoted without its run count is not a measurement, it is a headline. If a tool reports that a brand appears in a given share of AI answers, the first question is how many answers, and the second is whether those runs came from one session or several days. Ours came from single sessions, which is exactly why we are not claiming these percentages describe August; they describe these runs.
We have measured the day-to-day version of this separately, and it is worse: re-running the same question two days later reproduced the counts almost exactly while replacing most of the names. And measuring six categories rather than one, the number of runs a category needs to be described at all ran from about three to about thirty, which is why a single house standard for run count does not survive contact with a second vertical.
You do not get the runs you ask for, and that is not a footnote
Every question in this study asked for three runs. Ten of the nineteen came back with two or one, and the measurement said so each time rather than padding the total. When we went back on 10 August to add four more questions in France and Germany, three of the four stopped early and one returned a single run, which is why that one carries no percentage anywhere in this article.
That is worth publishing on its own. A stability figure is a fraction, and the denominator is not something you choose, it is something you are given. If a tool asks an assistant for ten runs, receives six, and reports “cited in 60% of runs” against a denominator of ten, the number is wrong in the direction that flatters nobody in particular and misleads everybody.
The check is simple and nobody offers it: ask your vendor what happens when a run fails. If the answer is that the requested count is what gets reported, the percentages have an invented denominator in them, and the size of the invention is unknowable from outside. Our own answer is in the register: the run count column is what came back.
We re-ran two of them a day later, and the best one fell apart
Everything above is measured inside single sessions, which the limits at the bottom of this article have always said is the weakness. So on 10 August we re-asked two of the 9 August questions, same market, same surfaces, same method.
| Question | 9 August | 10 August |
|---|---|---|
| US law firm | 13 of 20 stable, 65%, 2 runs | 7 of 24 stable, 29%, 2 runs |
| US ecommerce | 7 of 24 stable, 29%, 3 runs | 7 of 23 stable, 30%, 2 runs |
The most stable question in the entire study was the least reproducible one in it. Law firms went from 65% to 29% in a day, with the run count identical on both days, so this is not the arithmetic effect described above. It is the same question getting a different answer.
Ecommerce, an ordinary mid-range result, came back at almost exactly the same number. The run counts differ there, two against three, so it is not a clean comparison and we are not treating it as one. What can be said is that the unremarkable figure was the one that reproduced.
That inverts the instinct. A single high stability score looks like the strongest evidence on the page and is the one to trust least, because there are many more ways to be unstable than stable and an outlier has further to fall. If you are going to re-check one number in a client report before you send it, re-check the best one.
And sometimes it survives, which is why the advice is to check rather than to doubt. A travel booking question measured the same week, on the same three surfaces and the same two runs, came back 24 of 25 sources stable and then 24 of 26 in a second session, with the same twenty four domains both times. A high figure is not suspect for being high. Re-measuring is what separates a real core from a lucky session, and only one of those two is worth writing a content brief against.
One thing we cannot tell you, and it is our fault rather than the method’s. The 9 August rows recorded how many sources were stable, not which ones, so these two days can be compared on size and not on identity. Whether the seven stable law firm sources on the second day are seven of the original thirteen or seven different ones is not answerable from what we wrote down.
A third session, same day, answers the part we could not. We asked the law firm question once more on 10 August, in a new session. It stopped after a single run, so it produces no stability rate and we derive none from it. One run can still answer a different question: which sources came back.
All seven of the second session’s stable sources appeared again: reddit.com,
answeringlegal.com, natlawreview.com, exults.com, rebuttalpr.com,
mycase.com and lawyerist.com. Ten of that session’s seventeen one-offs also
reappeared, and three domains were new. The two cited sets overlap at 0.63 by
Jaccard.
So the percentage is more volatile than the sources it describes. A number that moved 36 points sits on top of a core that did not move at all. That is worth separating in a client report: “your stability score fell” and “the pages holding up the answer changed” are different sentences, and this week only the first one was true.
The reason is mechanical rather than mysterious. “Cited in every run” is a criterion with almost no tolerance at two or three runs: one source missing from one run drops out of the numerator entirely, and with a numerator in single digits that is a large percentage swing from a small change in the answer. The core survives; the statistic describing it does not.
The obvious fix does not work, and that is worth a paragraph. If “cited in every run” is too brittle, the natural replacement is the average citation rate, which tolerates a source missing one run instead of discarding it. Applied to these two sessions it looks much better behaved: 0.825 falling to 0.646, a 22% relative drop against 55% for the stable share.
That improvement is an illusion, and the algebra says so. With two runs a source
is cited in one run or both, so its rate is 0.5 or 1, and the average is exactly
0.5 + 0.5 × (stable share). It is the same number rescaled onto a range half
as wide. It cannot be less volatile in any way that means anything, and it
carries no information the first number did not.
At three runs there are three levels, a third, two thirds and one, and the average stops being a linear transformation of the stable share. So tolerance is bought with runs, not with formulas. A cleverer statistic over two runs is cosmetics, and we nearly published this one as a finding.
Same vertical, different market: the four clean comparisons
Most cross-market comparisons in this data are contaminated by the run count. Three are not, because the vertical and the run count both match:
| Vertical | Runs | Gap | ||
|---|---|---|---|---|
| Hotels | 2 | Spain 56% (9 of 16) | Germany 21% (5 of 24) | 35 points |
| Law firms | 3 | France 32% (7 of 22) | Spain 17% (4 of 23) | 15 points |
| Ecommerce | 3 | Spain 32% (9 of 28) | United States 29% (7 of 24) | 3 points |
| Clinics | 3 | Spain 23% (5 of 22) | Germany 21% (6 of 29) | 2 points |
The clinics row is the one we went back for. The first version of this study had three of these comparisons and two of them disagreed, which is not enough to tell a pattern from a coincidence, so we ran ecommerce and clinics in France and Germany specifically to add clean pairs. It bought one: the German clinics question is the only one of the four that returned all three runs.
Two of the four differ a lot and two barely differ at all. That is the honest summary and a more useful one than a clean rule would have been. The market is not a modifier you can apply to a number. Spanish hotels produced the second most stable answer in the whole study and Spanish law firms produced the least stable, in the same market, in the same language, on the same day.
If there is a pattern in there, this study cannot see it. A later measurement found one on a different question, by asking the identical string twice rather than counting inside one reading, and the two numbers are not in conflict because they measure different things. What it can say is that an agency running the same measurement for two clients in two countries should expect the reliability of the two reports to differ, and should not assume the difference is the client’s fault.
The amount cited moves less than what is cited, and Germany is the outlier
| Market | Questions | Average distinct sources per question |
|---|---|---|
| United States | 7 | 21.1 |
| Spain | 4 | 22.3 |
| France | 4 | 23.3 |
| Germany | 3 | 26.0 |
Five sources between the widest and the narrowest, on question counts of three to seven. No market builds its answer out of dramatically less material, and the largest market is the one that uses the least.
That matters for anyone planning market entry on the assumption that a smaller market is an easier answer to get into. On this evidence it is not thinner, it is slightly wider and less stable: Germany cites the most sources per question and holds the fewest of them across runs.
What to do with this if you report to clients
Quote the run count next to every percentage. It is the single change that makes an AI visibility report defensible, and almost nobody in this category does it. Our own numbers move nine points on that alone.
Report the stable set separately from the long tail. A client who sees twenty cited sources and is told they are absent from all twenty will react differently from one who is told the answer has a core of six and a tail of fourteen that changes every time. The second is true and the first is a scare.
Do not build a content plan against a one-off citation. Fifty eight percent of what you would be planning against will not be there when you check. The third that holds is where a brief is worth writing, and it is small enough to name page by page.
Ask a vendor whether their repeated runs bypass the cache. If they do not, a daily trend line is a flat line by construction, and a stability score built on it is measuring their infrastructure rather than your visibility. This is a question with a right answer and it is cheap to ask.
And measure the market you sell in, not the one your reporting tool defaults to. Every run here had country and language set explicitly, which is not the default anywhere, and it is the difference between four results and one result copied four times.
What this study does not show
- Nineteen questions is small. Four markets, eight verticals, and four verticals appear in more than one market. Every percentage here has a denominator in the tens, not the thousands.
- One question is excluded from every percentage. The German ecommerce question returned a single run, at which size every source is trivially both stable and once-only. It is in the register because the register records what happened, and out of the analysis because it cannot answer the question the analysis asks.
- Two runs and three runs are not the same experiment, which is why they are reported apart rather than pooled into one headline. The 32% aggregate mixes them and should be read as a midpoint, not a rate. It survived the study growing from fifteen questions to nineteen and from two markets to four, which is evidence that it is a stable midpoint rather than proof.
- Two questions have a second day, seventeen do not. Runs within one session tell you about that session, and the two we re-asked disagree about how much that matters. Claiming a source is durably cited needs days, not runs, across the whole set.
- Three surfaces, not four. ChatGPT is excluded because the route we use returns no source list for it, and reporting that as zero would be reporting our instrument.
- One phrasing per question. A Spanish study of this same class of question used four phrasings and found 49 sources cited with not one appearing in all four, so these sets belong to these phrasings and not to the intent behind them.
- Citation is not endorsement, and absence is not a verdict. A cited source answered the question well. An absent one may be excellent and simply unwritten about.
- Nothing here is causal. We changed nothing and recorded what came back.
- The German hotel question is ragged. AI Overview answered once where the other two surfaces answered twice, and it is recorded as two runs rather than smoothed.
Common Questions About This Study
How stable are AI search results for agency buying questions?
About a third of cited sources held across every run, and 58% appeared exactly once. Across nineteen questions in four markets, no single question had a stable majority: the best was 13 of 20 sources over two runs and the worst was 4 of 23 over three. Any single AI answer to a buying question is mostly made of sources that will not be in the next one.
Why does the number of runs change an AI visibility score?
Because “cited in every run” is an easier bar over two runs than over three. In this study, questions that ran twice showed 37% of sources stable and questions that ran three times showed 27%, measured the same week with the same method. Ten points of an apparent stability score are the run count and nothing else, which is why a percentage published without its denominator cannot be compared to anything.
Do AI answers differ between countries for the same question?
Sometimes by a lot and sometimes not at all. Comparing only questions with the same vertical and the same run count, Spanish and German hotel answers differed by 35 points of stability, French and Spanish law firm answers by 15, Spanish and American ecommerce answers by 3, and Spanish and German clinic answers by 2. Two large gaps and two small ones, from four comparisons, is not a rule about markets.
Should an agency report AI visibility to clients?
Yes, with two things attached: the number of runs behind every percentage, and a split between the sources that appear every time and the ones that appear once. Without the run count the figure cannot be compared to the next report. Without the split, a client reads a long tail of coincidences as a competitive position.
Which AI surfaces were measured?
Google AI Overview, Google AI Mode and Gemini, pooled per question, with the cache bypassed so every run is a fresh call. ChatGPT is deliberately absent: the route used here returns no source list for it, and a zero produced by an instrument is indistinguishable from a real zero.
How can I repeat this measurement?
Ask the same question with the country and language set explicitly, at least three times, on each surface separately, with caching off, and record what actually ran rather than what you requested. That last part is not pedantry: ten of our nineteen questions came back with fewer runs than we asked for. Then count three things: distinct sources, sources present in every run, and sources present once. Those three numbers are the whole study, and the third one is the one nobody publishes.
Where this leaves you
The finding we did not expect is how little of an AI answer is load-bearing. Nineteen questions, four markets, three surfaces, and 58% of everything cited was a one-off. That figure did not move when the study grew by four questions and two markets, which is the closest thing to a replication it has.
The part worth working on is the third that holds, and it is small enough to name: six or seven sources per question. That is a content brief, not a dashboard, and it is the same conclusion we reached measuring what AI cites for universities, where the stable core was student housing and tutoring firms rather than the institutions themselves.
If you run visibility reporting for clients, the cheapest improvement available to you is not a better tool. It is printing the run count next to the percentage, and separating the core from the tail. Both are free, and both are the difference between a report a client can act on and a number they will quote back to you when it moves.
We have also asked which agency the assistants name when the vertical is GEO itself, and the answer had the same shape.
If you would like that run for your own accounts, we do this for agencies.
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.