Skip to content
AI VisibilityMetrics
EN

Share of Voice in AI Answers: How It Is Actually Measured

Every tool sells share of voice and none publishes how many runs are behind it. How it is calculated, and how many runs it takes.

· Updated · 11 min read

Every tool in this category sells share of voice, and not one of them publishes how many runs the number is built on. That single missing figure decides whether the percentage on your dashboard is a measurement or a decoration.

So here is the arithmetic, and a measurement of the question itself.


The short version

  1. Share of voice is not visibility, and confusing them is the most common error in this category.
  2. A single run does not produce a share of voice. It produces a draw.
  3. At a hundred runs, 30% actually means “between 21% and 39%”. That is the best sampling anyone in this market publishes.
  4. It is calculated per market or it is not calculated. Averaging countries hides being invisible in three of four.
  5. Some categories have no number at all, not zero, because there is no answer to measure.

Share of voice and visibility are not the same metric

Visibility asks how often you appear at all: in what share of tracked prompts is your brand named, on any terms, in any position.

Share of voice asks how the naming is divided: of every brand mention in your category’s answers, what fraction is yours. One is presence. The other is competitive.

A brand can hold high visibility and low share of voice, which means it gets named in most answers and is always the fourth name in a list of six. The opposite happens too, in narrow categories where only two brands are ever mentioned and being one of them is worth more than a broad presence would be.

Tools that report one and label it the other are not rare. Ask which is on the chart.

Why one run does not produce a share of voice

Here is the flaw that makes most reporting in this category useless.

AI answers vary between identical requests, and less than we used to say. We asked Google AI Overview the same question three times in Spain, with country and language fixed: eleven domains were cited, five appeared in all three runs and four appeared once. An earlier version of this article reported seventeen domains with fourteen appearing once, which came from a count that also swept in the searches behind the answer rather than only its citations.

If the set of cited sources changes between runs, so does the set of brands named. A share of voice computed from one answer is a photograph of a draw, not a rate.

That is not an opinion, it is arithmetic. And it has an uncomfortable consequence for nearly every dashboard in this market: if your tool fires each prompt once a day and shows you a percentage with a decimal point, that decimal is decoration.

How many runs it actually takes

Nobody publishes this, so here it is. A share of voice is a proportion, and the uncertainty of a proportion has had a formula for a century.

For a rate near 30%, this is what each sample size buys you, as a 95% confidence interval:

Runs per promptWhat you can claim
1Nothing. You observed a yes or a no
1030%, plus or minus 28 points
3030%, plus or minus 16 points
10030%, plus or minus 9 points
40030%, plus or minus 4.5 points

Look at the hundred row, because it is the interesting one. A hundred runs per prompt per model is the best declared sampling in this market, Evertune’s, and even that only tells you your share of voice is somewhere between 21% and 39%.

It is an enormous improvement on one run, which tells you nothing. It is also not the precision the dashboards imply.

Now the hard number. Reliably detecting that your share of voice moved from 30% to 35% takes roughly 1,400 runs per condition, at 95% confidence and 80% power. Not a hundred. Not monthly. Per condition and per prompt.

Which produces the advice that actually helps:

  • Stop chasing small movements. Three points of variation on a panel built from a hundred runs is noise with a number attached.
  • Aggregate across the prompt set. Fifty prompts at a hundred runs is five thousand observations of your brand. The per-prompt rate stays noisy while the set-level rate becomes usable.
  • Widen the window. You cannot afford 1,400 runs a day. You can compare one fortnight against the next.
  • Use a control. A comparable competitor, measured the same way, absorbs platform drift, which is the largest source of movement nobody is measuring.

The arithmetic above uses the normal approximation to a binomial proportion at p = 0.30, and the power calculation assumes a two-sided test at α = 0.05. Both are textbook and both can be checked, which is more than can be said for most numbers in this market.

What the engines answered when we asked this exact question

On 6 August 2026 we put the title of this article to four surfaces at once, in English, United States: ChatGPT, Gemini, Google AI Overview and Google AI Mode.

This is one run per surface. By the argument above, that makes it a draw and not a rate, and we are reporting it as a draw. The ChatGPT figure was first published as 28 and is corrected here to 6: the larger number mixed the sources the answer cited with the results its search returned, and the correction is explained in full. What one draw can still show is which domains were reachable enough to be pulled into an answer at all.

Domains in the response record
ChatGPT6
Gemini7
Google AI Overview10
Google AI Mode22

Four surfaces, and the number of domains behind the answer ranges from six to twenty-two for the same question asked the same way.

The distribution of the sources is the useful part:

DomainSurfaces citing it
semrush.com4 of 4
llmpulse.ai3 of 4
nightwatch.io3 of 4
shadow.inc3 of 4
cognizo.ai3 of 4
blog.hubspot.com2 of 4
netranks.ai, siftly.ai, dageno.ai, trygeometrics.com2 of 4 each

Two things fall out of that table.

The citation goes to the explainer, not to the vendor. Semrush was cited on all four surfaces, and the page doing the work is a blog post explaining how to measure AI share of voice. Not a pricing page, not a product page, not a feature list. The assistants were asked which tools measure a thing, and they went to the article that explains the thing.

The named set is not the catalogued set. The tools the four surfaces named between them were Profound, Semrush, Peec, Otterly, Scrunch, Writesonic, Conductor, Athena, LLM Pulse, Spotlight and HubSpot’s AEO product. Several of the most-cited domains belong to companies absent from our catalogue we built by reading vendor pricing pages one by one. Whatever decides which names an assistant produces, it is not the same thing that decides which vendors exist.

A share of voice with no market means nothing

Averaging across countries is the quiet way to make a bad number look acceptable.

Answers are geographic. The same question in Spanish from Madrid and in English from Chicago produces different brands, different sources and, frequently, different categories. We have documented a case where a single model release turned European brands into American ones inside 24 hours across three sectors and two countries.

A brand that holds 40% share of voice in one market and 0% in three others reports 10% when you average it. That number describes nothing that happened anywhere, and it hides the only fact worth acting on.

Measure per market, per language and per engine. Report them separately. If a tool cannot separate them, it is not measuring share of voice, it is measuring an average of unlike things.

When the number is not zero but undefined

There are categories where AI answers do not name brands at all, or where the question does not trigger a generative answer in the first place.

In those, a share of voice of 0% is misleading, because zero implies a competition you lost. The honest report is that no answer exists to be measured, which is a different finding and often a better one: it means the surface has not opened yet, and being early is cheaper than being late.

Any tool that shows you 0% without telling you whether an answer was produced at all is giving you a number that cannot be acted on.

How to calculate it yourself

  1. Freeze a prompt set that reflects how buyers actually ask, and do not change it. A changed prompt set makes every comparison against history invalid.
  2. Fix the market: country and language, declared, per run.
  3. Run each prompt many times. Not once. Whatever budget you have, spend it on repetitions before you spend it on more prompts.
  4. Count two things separately: whether your brand was named, and whether your own site was cited as a source. They move independently and mixing them makes any analysis ambiguous.
  5. Divide by the category total, not by the run count. Share of voice is your mentions over all brand mentions.
  6. Report an interval, not a point. With n runs at rate p, the 95% interval is roughly p ± 1.96 × √(p(1−p)/n).
  7. Measure a control brand the same way, so platform drift does not read as your result.

What this article does not prove

  • The four-surface measurement above is one run each. It shows what was reachable in that draw. It is not a rate, and we have not repeated it enough times to make it one.
  • The three-run and five-run figures are small samples from our own testing, quoted as illustrations of variance, not as measurements of it.
  • The sample-size table is arithmetic, not empirical. It tells you what a proportion’s uncertainty is at a given n. It does not tell you that any particular tool’s numbers are wrong, only what they would have to publish for you to know.
  • We have not audited any vendor’s actual run count, because none of them publishes it. That is the finding.

Common Questions About AI Share of Voice

What is share of voice in AI answers?

It is your brand’s share of all brand mentions in the answers a generative engine produces for your category. If ten brands are named across your tracked prompts and three of the mentions are yours, your share of voice is 30%. It is a competitive metric, unlike visibility, which only asks whether you appear at all.

How is it different from share of model?

Share of voice measures answers produced with live retrieval, where the model searches and cites. Share of model measures what the model says from its training weights with no search. They can differ enormously for the same brand, and a tool that does not tell you which one it is running is not telling you what it measured.

How many runs do I need for the number to mean anything?

For a rate near 30%, a hundred runs per prompt gives you plus or minus 9 points at 95% confidence. Ten runs gives you plus or minus 28, which is not a measurement. Detecting a five-point change reliably takes around 1,400 runs per condition, which is why the useful comparison is fortnight against fortnight across a whole prompt set, not day against day on one prompt.

Can I measure share of voice across several countries at once?

You can, and the result will not describe any of them. Answers are geographic: the same question produces different brands and different sources per market. A brand at 40% in one country and 0% in three reports 10% on an average, which is true arithmetic about nothing. Measure and report per market.

What if my category has no AI answers?

Then your share of voice is undefined, not zero. Zero implies a competition you lost. Undefined means the surface has not opened for your category yet, which is useful information and usually good news, because arriving before the answers stabilise is cheaper than arriving after.

What should I ask a vendor before buying?

Four questions. How many times does each prompt run, per model, per day. Do you report an interval or a point. Do you separate brand mentions from citations of my own site. And can you break the number down by market, language and engine. Only one tool in our comparison of every tool we verified publishes an answer to the first one.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)