Skip to content
AI VisibilityGEO
EN

We Asked Fifteen Frozen Questions Again Two Days Later. The Best Score in the Set Did Not Survive Its Own Control.

Thirteen categories re-asked after two days. Median source overlap 0.63, and the best and worst rows both moved when read again the same afternoon.

· Updated · 15 min read

Every visibility tool sells a line on a chart, and a line needs the same question asked on two different days. We had never done it. On 10 August we froze fifteen buying questions and recorded, by name, every domain that came back in all three runs. On 12 August we asked the same fifteen on the same surface, in the same markets, with the same three runs. Thirteen rows are comparable, and their median overlap with the earlier set is 0.63.

That number is almost exactly what two sessions on the same afternoon already cost us. The interesting part is not the median. It is that the highest and lowest rows in the set both moved when we simply read them again.

Disclosure: EchoWi sells AI visibility measurement, so a study arguing that day over day movement is mostly noise is a study with an interest, and it cuts at a chart we could sell too. The fifteen questions, the surface, the two markets, the run count and both dates are stated below, and every row including the two controls is in our measurement register, so anyone can run it against us.


The short version

  1. Thirteen of fifteen rows are comparable, meaning three answered runs on both days. Median overlap of the durable sets is 0.63, and the range is 0.00 to 1.00.
  2. That median is not distinguishable from asking twice in one day. Two sessions on the same day have run 0.67 to 0.70 in our earlier work, and repeats inside a single session 0.63. Two days landed on the same ground.
  3. Both controls moved, and that is the finding. The one row that returned an identical set after two days fell to 0.67 when re-read that afternoon. The one row that returned nothing durable rose to 0.25.
  4. A flawless daily stability score predicts nothing about two days later. Four categories scored 100 per cent on both readings. Their cross-day overlap ran from 0.25 to 0.75.
  5. What persists is a publisher layer, not a category’s own sources. YouTube held in ten of the thirteen categories, Reddit in four. Across the nine United States rows, 3 of 19 vendor domains survived two days against 32 of 59 for everything else.

What was frozen, and why that is the whole trick

A stability row usually records a count: this question cited twenty-two domains and seven of them came back every run. That is enough to publish a percentage and not enough to ask this question. Two days later you can see whether the count moved. You cannot see whether the sources did, and those are different claims that a count cannot separate.

The 10 August rows recorded the durable domains by name. That is the only reason this study exists. Our own travel booking rows from the same week recorded counts alone, so they can never answer it, and that is a design fault in our register rather than a limitation of the surface.

So the frozen set is fifteen high intent buying questions, nine in the United States in English and six in Spain in Spanish, each asked on one surface with three runs. On 12 August the identical strings went out again with identical settings. Two rows dropped out of the comparison and are named rather than dropped quietly: the United States help desk question returned no summary at all on the second date, which is a different event from citing nothing, and the Spanish password question answered once of three, so its durable set is a single run set and is not comparable to a three run one.

The thirteen rows

Overlap is the shared domains divided by the union of both sets. Kept counts domains present in all three runs on both dates.

MarketCategoryDurable 10 AugDurable 12 AugKeptOverlap
USweb hosting5551.00
USonline stores4540.80
EShelp desk4540.80
USCRM7760.75
ESonline stores3430.75
ESaccounting7550.71
USpassword managers7650.63
USbusiness banking4430.60
USproject management4530.50
USpayroll11650.42
USaccounting13740.25
ESweb hosting3310.20
USemail marketing23000.00

The median row is United States password managers, which shared five domains out of a union of eight. Anyone can repeat that division.

The two controls, and they are the finding rather than a footnote

A table like the one above invites you to read the top row as a stable category and the bottom row as a volatile one. Both readings are wrong, and the cheapest way to find that out is to ask the same two questions again the same afternoon.

The best row. Web hosting in the United States returned the identical five domains two days apart: Reddit, CNET, YouTube, Forbes and TechnologyAdvice. A perfect 1.00, the only one in the set. Read again a few minutes later, with the same three runs, CNET leaves and ZDNet arrives. Four shared of a union of six is 0.67, so the same row measured twice on the same day produces two different answers about what two days did.

The worst row. Email marketing in the United States returned nothing at all that survived all three runs, which scores 0.00 against a baseline of twenty-three durable domains. Read again the same afternoon it returns seven, and those seven are exactly the seven that had reached two runs of three an hour earlier. Recomputed against the same baseline, the overlap is 0.25.

The second control is the one that shows the machinery. The sources were present in both readings. What changed was whether every run caught them, and a durable set has no tolerance at three runs: one distracted run drops a domain out of the numerator entirely. A 0.00 there did not mean the category had emptied out. It meant the bar is all or nothing and the coin landed badly.

So the honest reading of the table is that no single row is a finding, in either direction. The median across thirteen rows is the citable number, with both controls printed beside it, and even that is one draw.

A perfect daily score predicts nothing about the week

Four categories returned a durable share of 100 per cent on both dates, meaning every domain the answer cited came back in all three runs, twice, two days apart. On any dashboard that is a category with maximally stable citations.

CategoryDaily stability, both daysCross-day overlapKept
CRM100%0.756 of 7
business banking100%0.603 of 4
payroll100%0.425 of 11
accounting100%0.254 of 13

Accounting scored flawlessly on both readings and replaced nine of its thirteen sources in between. QuickBooks, Xero, Wave, FreshBooks, Invoicefly and Lettuce were all durable on 10 August and none of them was durable on 12 August. The daily number never noticed, because it can only see inside one reading.

Across the thirteen rows, the correlation between the daily share on the first date and the cross-day overlap is negative 0.30. Not a strong relationship in either direction, which is the point: the number your dashboard shows today carries essentially no information about whether the same sources will be there on Thursday.

What persists is a publisher layer, not the slots a vendor can win

Splitting the durable domains by what they sell changes the picture, and this is the part with a commercial edge.

Vendor domains keptEverything else kept
United States, 9 rows3 of 1932 of 59
Spain, 4 rows7 of 96 of 8

In the United States rows the sources that sell the product being asked about are the ones that churn, while publishers, forums and video hold. In the Spanish rows there is no gap at all. That is a real split and it is on small numbers, nine vendor observations in Spain against nineteen in the United States, so it goes here as a recount and not as a rate, and it needs its own study before anyone calls it a market difference. That study now exists: asking the same three categories in all four markets shows most of this gap was category mix, because accounting is the one category that gives vendors slots anywhere.

What crosses categories is a short list. Counting how many of the thirteen categories kept each domain across both dates: YouTube in ten, Reddit in four, Forbes in three, Zapier in three, PCMag in two, CNBC in two. Those six are the only domains that held in more than one category, and none of them sells CRM, payroll or hosting.

That is the durable structure the question is actually asking about. A category’s own vendor list turns over. The general purpose layer above it does not.

What to do with this if you are buying a tracker

Three questions, in order of how much they change what you are paying for.

Does a movement in your chart come with the domains? A percentage that fell from 62 to 41 could mean the sources changed or could mean one run in three missed two of them. Those are opposite situations and only the list of names tells them apart. If the product cannot show you which domains entered and left, the line is not reportable.

How many phrasings does it follow per intent? Our earlier work found four wordings of the same Spanish hotel question sharing zero domains that held across all four. A tracker following one string is following the stability of a sentence.

What is the run count behind a single point? At three runs, a durable set is all or nothing, and the email marketing control shows what that costs. Tolerance is bought with runs, not with a cleverer statistic.

Four points on the axis, and the drift does not clear its own threshold

This piece measured two days. The same five United States cells have since been re-read at four days and at six, always against the same 10 August baseline, so the axis now has four points rather than a before and an after.

Distance from the baselineComparable cellsMedian overlap
Same session, same day0.67 to 0.70
Two days90.60
Four days50.50
Six days40.32

The two-day figure here is 0.60 and the headline of this article is 0.63, because they are different sets: 0.63 is every comparable row in every market and 0.60 is the United States alone, which is the only market the later re-reads cover. Comparing a two-market number with a one-market one would make the first step of the axis look steeper than it is.

Six days is below every earlier point, and it does not clear the bar we set before running it. The threshold written down first was that a drift claim needs the median to fall and four or five of the five cells to fall against their own four-day reading. The median fell, from 0.50 to 0.32 with a range of 0.24 to 0.42. The cells went 3 down and 1 up out of 4, with 1 excluded because web hosting stopped at two runs and then at one, and a two-run durable set is a different bar.

So this is published as what it is: a fourth point that continues a decline, on five fewer comparable cells than the two-day arm, with a per-cell direction that falls short of what we said would convince us. Anyone quoting 0.32 as the six-day drift rate is quoting a median of four numbers.

What is not happening is the answers getting narrower. Accounting went from 13 durable domains to 11, payroll from 11 to 15, password managers from 7 to 10. The overlap falls while the sets hold their size or grow, so this is substitution rather than shrinkage: different sources, not fewer.


What this does not show

One surface and two markets. Everything here is a single answer surface, in the United States in English and Spain in Spanish. Nothing in it transfers to another surface or another country without measuring there.

Fifteen questions, one baseline reading, one repeat. The 10 August figures are a single reading, and so is each 12 August figure. The two controls exist precisely because we cannot separate a two day effect from ordinary reading to reading movement at this sample size, and they suggest the two are the same size.

Two days is not a week or a quarter. A gap this short is the cheapest possible version of this study. It says a two day gap is not distinguishable from an afternoon. It says nothing about whether a real seasonal shift would show up, and we would expect one to.

The all or nothing bar is doing work. A domain in two runs of three is invisible to a durable set. That choice is defensible and it is a choice, and the email marketing row is what it costs when it lands badly.

Two of fifteen rows are not comparable, and both are named above rather than smoothed into the denominator.

Nothing here is causal. These are observations of a surface that re-plans retrieval per request. No intervention was made and none of these movements has an established cause.

For how a single day’s percentage behaves and why breadth drives it, the companion piece on what a stability score actually measures pools ninety one of our published rows. For the frozen set this study repeats, see who holds the durable citation slots.

Common Questions About Tracking Citations Over Time

How much do AI citations change from one day to the next?

Across thirteen buying questions asked on one answer surface two days apart, the median overlap between the sets of domains that held in every run was 0.63, on a scale where 1.00 means an identical set. The range across categories was 0.00 to 1.00. That median is level with what we measure between two readings taken the same afternoon, so at this sample size a two day gap does not cost measurably more than asking again after lunch.

Does a stable citation score mean my sources are not changing?

No, and this is the most expensive misreading available. Four categories scored a perfect daily stability figure on both measurement dates, meaning every cited domain came back in every run, twice. Their two day source overlap still ran from 0.25 to 0.75. One category kept four of its thirteen sources while scoring flawlessly on both days. A daily score can only see inside a single reading.

Why did one question return no stable sources at all?

Because a durable set counted at three runs is all or nothing: a domain missing from one run drops out completely. The email marketing question returned nothing durable in one reading and seven durable an hour later, and those seven were exactly the domains that had reached two runs of three in the first reading. The sources were present both times. Only whether every run caught them changed.

What should I track instead of a stability percentage?

Track the set of domains, and track it across several categories rather than one. In this study the only things that held across more than one category were general purpose sources: YouTube in ten of thirteen categories, Reddit in four, Forbes and Zapier in three each. A vendor’s own presence in its category is the volatile part, which is inconvenient because it is also the part you are paying to move.

How many runs does a citation figure need?

More than three if you want a durable set with any tolerance, because at three runs the bar has none. The alternative of switching to an average citation rate does not fix it: with two runs that average is a linear rescaling of the same share and adds no information. Tolerance is bought with runs, not with a different formula.

Can I repeat this measurement myself?

Yes, and that is the point of publishing the questions. Freeze a set of buying questions in your category, record for each one the domains that appear in every run rather than just the count, and ask the same strings again on the same surface after a gap. Recording the names rather than the count is the step that makes the comparison possible at all.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)