We Asked for the Same Hotel Booking Four Ways. In Spain the Number Barely Moved and Not One Source Survived.
One intent, four wordings, two markets. In Spain the stable share held between 30% and 48% while zero domains survived all four phrasings.
We took one question, how to book a hotel, and asked it four different ways in the United States and four different ways in Spain. In Spain the stable share came back at 42%, 41%, 48% and 30%, which looks like a reliable measurement of a category. Across those four wordings, the number of sources that survived all four was zero.
That is the finding, and it is worse than instability. A metric that moves tells you it cannot be trusted. This one holds still while the thing underneath it is replaced completely.
Disclosure: EchoWi sells AI visibility measurement and this study contradicts something we published five days earlier. The answer to a conflict of interest is a method you can check, so the four wordings, the two markets, the three surfaces, the run counts and the date are all below, and every row is in our public measurement register.
The short version
- In Spain the rate held and the sources did not. Four wordings returned 42%, 41%, 48% and 30% stable. Across all four, zero domains appeared in every run of every wording.
- In the United States rewording moved the rate by 57 points, from 92% down to 35%, with nothing changed but the words.
- Our own best published result does not survive this. The travel study found 24 domains holding across runs and called it a core you could build against. Ask the same thing four ways and the core is three domains:
reddit.com,youtube.comandfrommers.com. - Phrasing sensitivity is itself a property of the market. The American question swung 57 points across wordings. The Spanish one swung 18.
- One wording could not reliably produce an AI Overview at all, in either of two sessions, which is a fact about that sentence and not about the surface.
What we measured
| Intent | Booking a hotel, held constant |
|---|---|
| Wordings | 4 per market, in the market’s own language, each a way a person actually asks |
| Markets | United States and English, Spain and Spanish, each set explicitly |
| Surfaces | Google AI Overview, Google AI Mode and Gemini, pooled per wording |
| Runs | 2 per wording, recorded per wording rather than assumed |
| Cache | Bypassed. Every run is a fresh upstream call |
| Date | 10 August 2026 |
The four wordings, in English, were “what is the best website to book hotels”, “where should I book a hotel online”, “which hotel booking site is most reliable” and “best site for cheap hotel deals”. The Spanish four mirror them. Every one of the eight is in the register with its prompt exactly as it was sent.
Two rows carry a note rather than a clean reading, and both are in the register. The second English wording was measured twice because its first session looked broken; it was not, and both sessions are reported. The third Spanish wording stopped after a single run on its first attempt, which makes every source trivially stable, so that attempt is excluded from every figure here and kept in the register because it happened.
This design has one purpose: to attack our own strongest published claim. Five days ago we reported an American hotel question holding 24 of 25 sources across runs, then 24 of 26 in a second session, the same twenty four domains by name, and we wrote that six or seven stable sources is a content brief and twenty four is a map. This study asks what that map is a map of.
The eight cells
| Market | Wording | Stable share | Cited | Held |
|---|---|---|---|---|
| United States | “what is the best website to book hotels” | 92% | 24 | 22 |
| United States | “which hotel booking site is most reliable” | 83% | 18 | 15 |
| United States | “best site for cheap hotel deals” | 58% | 31 | 18 |
| United States | “where should I book a hotel online” | 35% | 20 | 7 |
| Spain | “¿qué web de reservas de hotel es más fiable?” | 48% | 25 | 12 |
| Spain | “¿cuál es la mejor web para reservar hoteles?” | 42% | 31 | 13 |
| Spain | “¿dónde reservo un hotel por internet?” | 41% | 27 | 11 |
| Spain | “mejor web para ofertas de hotel baratas” | 30% | 20 | 6 |
| cited in every run | distinct sources cited | |
|---|---|---|
| US · best website | 22 | 24 |
| US · most reliable | 15 | 18 |
| US · cheap deals | 18 | 31 |
| US · where should I | 7 | 20 |
| ES · más fiable | 12 | 25 |
| ES · mejor web | 13 | 31 |
| ES · dónde reservo | 11 | 27 |
| ES · ofertas baratas | 6 | 20 |
Read the two blocks differently, because they fail in two different ways. The American block is a 57 point drop that any dashboard would show as a problem. The Spanish block is four numbers inside an 18 point band, which any dashboard would show as a category behaving consistently.
The Spanish block is the dangerous one
Four wordings, 42%, 41%, 48%, 30%. If you measured one of them weekly you would report a stable figure around 40% and nobody would question it.
Now count the sources instead of the percentage.
| Domains that held | Shared with the wording above | |
|---|---|---|
| “¿cuál es la mejor web para reservar hoteles?” | 13 | |
| “¿dónde reservo un hotel por internet?” | 11 | 4 |
| “¿qué web de reservas de hotel es más fiable?” | 12 | 4 |
| “mejor web para ofertas de hotel baratas” | 6 | 2 |
Across all four Spanish wordings, the number of domains that held in every run of every wording is zero. Sixty one distinct domains appeared somewhere in the eight Spanish runs. Not one of them was reliably there for all four ways of asking.
The pairwise overlaps say the same thing more precisely. Measured as intersection over union, the four Spanish wordings share between 0.06 and 0.32 of their stable sets. The two closest, at 0.32, still disagree about two thirds of what they cite.
So the honest description of the Spanish result is: the rate is a property of the category and the sources are a property of the sentence. Those are two different facts and only the second one tells you which page to write.
The American block, and what it costs us
The American wordings drop from 92% to 35%, and the tempting read is that some wordings are simply worse. That is not what the sources say either.
Across the four American wordings, 53 distinct domains appeared. Three of them held in every run of every wording: reddit.com, youtube.com and frommers.com. Pairwise, the stable sets share between 0.16 and 0.38.
Five days ago we published that the American hotel question returned the same twenty four domains in two separate sessions and that twenty four stable sources is a map you can build against. That measurement was correct and we would run it again. What this study shows is that the map is a map of one sentence. Change the sentence to another one a customer would plausibly type, and twenty one of those twenty four stop being reliable.
We are not retracting the earlier figure. We are bounding it, which is the thing a vendor selling this category almost never does to its own best number.
Phrasing sensitivity is not a constant either
The two markets do not just differ in level, they differ in how much the wording matters at all.
| Spread across four wordings | |
|---|---|
| United States | 57 points, from 92% to 35% |
| Spain | 18 points, from 48% to 30% |
That is the same shape the travel study found when it varied the market instead of the words: two questions three points apart in Germany and 53 apart in the United States. Both studies land in the same place from opposite directions. Neither the market, nor the vertical, nor the question, nor the wording is a coefficient you can apply. Each combination is its own measurement, and the only way to know a cell is to measure that cell.
There is also one practical detail worth more than it looks. The wording “where should I book a hotel online” produced an AI Overview in only one of two runs, in both sessions we ran. Some ways of asking do not reliably trigger that surface at all, and a report that pools surfaces will show that as low visibility rather than as absent inventory.
What this study does not show
- One intent, two markets, eight comparable cells. Every percentage has a denominator in the twenties or thirties.
- Two runs, not three. “Cited in every run” is an easier bar over two runs, so these numbers are not comparable to three-run measurements elsewhere, including parts of our agency study, which measured that effect at about ten points.
- Four wordings is not the space of wordings. We chose four that a buyer would recognise. A different four would give different numbers, and that is the point rather than a defect. It is also worth saying that a growing share of the sentences reaching a site are not typed by a buyer at all but assembled by software, which is how a competitor’s system prompt turned up in our own Search Console.
- Three surfaces, not four. ChatGPT is excluded because the route we use returns no source list for it, and reporting that as a zero would be reporting our instrument.
- One session per cell, except two. The second English wording has two sessions and the third Spanish wording has a discarded first attempt. Both are named above and both are in the register.
- We did not test whether the wordings are equally common. Search volume would tell you which sentence to optimise for; this study only tells you that the choice matters.
- Nothing here is causal. We changed the words and recorded what came back, with no intervention on any page.
Common Questions About This Study
Does rewording a question change which sources an AI cites?
Yes, and by more than changing the country does. Asking about hotel booking four ways in the United States moved the share of sources cited in every run from 92% to 35%, and only 3 domains out of 53 held across all four wordings. In Spain, 0 domains out of 61 held across all four.
If the percentage stays the same, is my AI visibility stable?
Not necessarily, and this is the trap the study was built to show. Four Spanish wordings returned 42%, 41%, 48% and 30%, a band narrow enough to look like a reliable measurement, while sharing no domain at all across all four. A rate holding still is not evidence that the sources holding it are the same sources.
How many prompts should I track per question?
More than one wording, and the number depends on the market. On the same intent, the American wordings spread 57 points and the Spanish ones 18, so a single prompt is a much worse proxy in the United States than in Spain. There is no universal number, which is why we publish the cells instead of a rule.
Does this contradict your earlier travel study?
It bounds it. That study reported an American hotel question holding 24 of 25 sources and repeating at 24 of 26 with the same domains, and that measurement stands. This one shows those twenty four are a property of that exact sentence: reword it and three survive. Both results are in the register and the earlier figure has not been changed.
Which AI surfaces were measured?
Google AI Overview, Google AI Mode and Gemini, pooled per wording, with the cache bypassed so every run is a fresh call. ChatGPT is deliberately absent: the route used here returns no source list for it, and a zero produced by an instrument is indistinguishable from a real zero.
Where can I see the raw rows?
In our measurement register, which carries all ten rows of this study with each prompt exactly as it was sent, its market, its surfaces, its run count and its date, including the two rows excluded or flagged and the reason for each.
Where this leaves you
If you report AI visibility to anyone, the instruction from this study is short. Track the sentence, not the topic, and track more than one sentence per topic. A single prompt gives you a number that is real for that prompt and tells you very little about the next one, even when the next one means the same thing to a human reader.
And if a tool shows you a stable percentage week after week, ask it which domains that percentage is made of. In our Spanish block the percentage would have looked fine for a month. Underneath it, four different ways of asking one question shared not a single reliable source between them.
We published the twenty four domain result five days ago and we still believe it. This is what it is worth, measured against the only test that mattered: asking the same thing in words we had not used.
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.