Skip to content
AI VisibilityResearch
EN

Two Travel Questions, Four Markets. Indistinguishable in Germany, 53 Points Apart in the US.

Hotels and flights asked in four markets on three AI surfaces. The gap between the two questions runs from 3 points in Germany to 53 in the United States.

· Updated · 11 min read

We asked two travel booking questions, hotels and flights, in four markets on three AI surfaces. In Germany we could not tell the two questions apart, 39% and 42%. In the United States they are 53 points apart. Same vertical, same week, same method.

That is the number to take into a client conversation about AI visibility: how stable your answers are is not a property of your industry. It is a property of the exact question, in the exact market, and the two do not combine in any way you can predict from either one.

Disclosure: EchoWi sells AI visibility measurement. That is a conflict, and the answer to a conflict is a method you can check rather than a promise to be neutral. The questions, surfaces, markets, run counts and dates are all below, and every row is in our public measurement register.

Updated 10 August 2026. After publishing this we read one question five times in a day and watched its stable share run from 39% to 89% while the number of durable sources stayed at seven or eight. That does not touch the 53 point American gap or the 45 point Spanish one, both of which are far outside that range and one of which replicated in a second session. It does retire one claim: a three point gap is not a finding. The German row is written here as two questions this instrument cannot tell apart, which is what it always was, and the five readings are in the follow-up study.


The short version

  1. The gap between two questions in one market runs from indistinguishable to 53 points. Germany’s hotels and flights come back at 39% and 42%, three points, which is inside what a re-reading moves. The United States’ come back at 96% and 43%.
  2. One question is genuinely, repeatedly stable. The US hotel question held 24 of 25 sources across two runs, then 24 of 26 in a second session, and the twenty four are the same twenty four domains.
  3. Nothing about “travel” explains any of it. Across the seven comparable questions the stable share runs from 14% to 96%.
  4. Travel is more stable than agency buying overall, and that hides more than it says. 47% of cited sources held here against 32% in our nineteen-question agency study, and the spread inside travel is far wider than the difference between the two verticals.
  5. The unstable end is worse than anything we have published. Spanish flights returned 35 sources with 5 holding, and 30 appearing exactly once.

What we measured

Questions2 intents, “what is the best website to book hotels” and “…to book flights”, in the market’s own language
MarketsUnited States and English, Spain and Spanish, France and French, Germany and German, each set explicitly
SurfacesGoogle AI Overview, Google AI Mode and Gemini, pooled per question
Runs2 per question, recorded per question rather than assumed
CacheBypassed. Every run is a fresh upstream call
Date10 August 2026

Seven questions carry a stability figure. Three more were measured and are excluded, each for a stated reason: the French flight question returned a single run, at which size every source is trivially stable and trivially unique; and an earlier attempt at the Spanish hotel question completed on two surfaces instead of three, so it draws from a smaller pool and cannot sit beside the rest. Both are in the register because the register records what happened.

The method matches the agency study deliberately, so the two can be compared without an argument about instruments.


The two questions, market by market

MarketHotelsFlightsGap
United States96% (24 of 25)43% (9 of 21)53 points
Spain59% (13 of 22)14% (5 of 35)45 points
Germany39% (7 of 18)42% (11 of 26)3 points
France50% (10 of 20)one run, no rate
cited in every rundistinct sources citedUS · hotels24 → 25Spain · hotels13 → 22France · hotels10 → 20Germany · flights11 → 26US · flights9 → 21Germany · hotels7 → 18Spain · flights5 → 35
Each row is one question in one market. The left dot is what held across both runs, the right dot everything cited at all. Google AI Overview, AI Mode and Gemini pooled, cache bypassed, 10 August 2026.
Stable sources against all cited sources for the hotel and flight questions in four markets
cited in every rundistinct sources cited
US · hotels2425
Spain · hotels1322
France · hotels1020
Germany · flights1126
US · flights921
Germany · hotels718
Spain · flights535

Read the German row and the American row together, because that is the finding. In Germany, asking about hotels and asking about flights produced answers three points apart, which this instrument cannot separate. That is not the same as saying they behave identically: it is saying one reading cannot tell them apart. In the United States the same two questions are 53 points apart, and a reader who measured one and assumed the other would be wrong by more than half the scale.

Spain sits between them in shape and at the bottom in level: 45 points apart, with the flight question returning the least stable answer in anything we have published.


The outlier held, which is not what happened last time

The United States hotel question came back with 24 of 25 sources cited in both runs. That is far above anything in our agency study, whose best result across nineteen questions was 13 of 20, so it got the treatment that study taught us: measure it again.

Sources citedStableShare
First session252496%
Second session262492%

And the twenty four are the same twenty four domains, checked by name and not by count: booking.com, kayak.com, expedia.com, reddit.com, youtube.com, google.com, forbes.com and seventeen more, in both sessions.

That is the opposite of what happened to the most stable question in the agency study, which fell from 65% to 29% overnight. Both results matter and they say different halves of one rule. A high number is not suspect for being high; it is worth the cheapest possible check, and sometimes the check confirms a core you can actually build against. Six or seven stable sources per question is a content brief. Twenty four is a map.


Why the German result is the useful one

The temptation with a table like this is to lead with 96%. The German row is worth more, because it is the one that tells you the shape of the problem.

If question and market interacted in a simple way, Germany would show the same gap the United States shows, smaller or larger but present. It shows three points, which is to say no resolvable gap at all. Whatever makes the American hotel question so much more stable than the American flight question is not a property of hotels, because it does not travel to Germany.

We cannot say what it is from this data, and this study is not going to pretend. What it can say is what an agency should do about it: measure the question you are being paid for, in the market you are being paid for. A stability figure from a neighbouring question in the same industry is not evidence about yours, and in one of our three markets it would have been wrong by 53 points.


Against the agency study, which used the same instrument

Agency buying questionsTravel booking questions
Questions197
Stable share32%47%
Cited exactly once58%53%
Range across questions17% to 65%14% to 96%

Travel looks more stable, and the aggregate is the least interesting line in the table. The range is the point. Travel’s stable share spans 82 points across seven questions; the agency set spans 48 across nineteen. A vertical average would have hidden both the 96% and the 14%, and those are the two rows a buyer would actually want to know about.

This is the same conclusion the agency study reached from the other direction, where two of four clean cross-market comparisons differed by 15 and 35 points and two differed by 3 and 2. Neither market nor vertical is a coefficient you can apply.


What this study does not show

  • Seven comparable questions is small. Two intents, four markets, and only three markets carry both intents. Every percentage has a denominator in the tens.
  • Two runs, not three. “Cited in every run” is an easier bar over two runs, so these figures are not comparable to three-run measurements elsewhere, including parts of our own agency study. That study measured the size of that effect at about ten points.
  • One session per question, except one. Only the US hotel question has a second session. The rest describe the session they were measured in, which is the limitation that turned into a finding the last time we tested it.
  • Three surfaces, not four. ChatGPT is excluded because the route we use returns no source list for it, and reporting that as a zero would be reporting our instrument.
  • The Spanish flight row is ragged. AI Overview answered once where the other two surfaces answered twice. It is recorded as two runs and flagged rather than smoothed, and it is the least reliable row in the table.
  • One phrasing per question. A Spanish study of a different category used four phrasings and found 49 sources with none appearing in all four, so these sets belong to these phrasings.
  • Nothing here is causal. We changed nothing and recorded what came back.

Common Questions About This Study

How stable are AI answers for travel booking questions?

It depends entirely on which question and which market. Across seven comparable questions the share of sources cited in every run ran from 14% to 96%. The United States hotel question held 24 of 25 sources and repeated that in a second session; the Spanish flight question held 5 of 35. There is no single figure for travel, and any vendor quoting one is averaging over a range that spans most of the scale.

Does the same question behave the same way in different countries?

No, and not even approximately. The hotel booking question returned 96% stable sources in the United States, 59% in Spain, 50% in France and 39% in Germany. That is a 57 point spread on one question with the market as the only variable, measured on the same day with the same surfaces. Which market you are in is not a detail of the same answer, it changes which sources exist: we measured the same brands losing most of their AI visibility the further a market sits from the United States.

On the aggregate, slightly: 47% of cited sources held across runs here against 32% in our nineteen-question agency study. The aggregate is misleading. Travel’s range runs from 14% to 96% and the agency study’s from 17% to 65%, so the spread inside travel is wider than the difference between the two categories.

Should I trust a high AI visibility score?

Trust it enough to check it once, which is cheap. Our agency study’s best result fell from 65% to 29% the next day. Our travel study’s best held at 96% and then 92%, with the same twenty four domains both times. The check is what separates them, and it is the only thing that does.

Which AI surfaces were measured?

Google AI Overview, Google AI Mode and Gemini, pooled per question, with the cache bypassed so every run is a fresh call. ChatGPT is deliberately absent: the route used here returns no source list for it, and a zero produced by an instrument is indistinguishable from a real zero.

Where can I see the raw rows?

In our measurement register, which carries every question in this study with its prompt as it was asked, its market, its surfaces, its run count and its date, including the three rows excluded from the analysis and the reason each one is out.


Where this leaves you

The question we started with was whether a vertical has a stability profile. It does not. Two questions inside one industry can be indistinguishable in one country and 53 points apart in another, and no amount of averaging over “travel” would have told you which country you were in.

For anyone reporting AI visibility, that collapses into one instruction: measure the question your client is actually judged on, in the market they actually sell in, more than once. Everything else on this page is evidence for why the cheap version of that does not work.

And if you find a number as high as 96%, check it again before you celebrate. Ours survived. The best number in our agency study did not, and the difference between those two outcomes was one repeat measurement.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)