Skip to content
AI VisibilityResearch
EN

We Asked Where to Study Four Ways. Four Sites Survived Every Wording, and Not One of Them Is a University.

Rewording a hotel question in Spain left zero stable sources. Rewording a university question in the same country left four, and none of them is a university.

· 13 min read

This morning we published that asking about hotel booking four different ways in Spain left zero sources cited reliably across all four. We ran the same design on a different question in the same country, choosing a university in Madrid, and four domains survived every wording. Not one of the four is a university.

Both results are real and they answer different questions. The first says a stable core can be an illusion produced by one sentence. The second says it is not always an illusion, and tells you who gets to own it.

Disclosure: EchoWi sells AI visibility measurement. That is a conflict, and the answer to a conflict is a method you can check. The eight wordings, the two markets, the three surfaces, the run counts and the date are all below, and every row is in our public measurement register.


The short version

  1. The collapse is not universal. In Spain, four wordings of a hotel question shared no reliably cited source. Four wordings of a university question shared four.
  2. The four survivors are all intermediaries: graddus.com, uniscopio.com, smartresidences.es and maclapiovera.es. Two comparison sites, a student housing operator and a local guide.
  3. No university appears in all four wordings, in either market. Institutions surface for particular sentences and vanish for others.
  4. France keeps almost nothing. One domain survived all four Paris wordings, relocation-in-paris.fr, and it is a relocation service.
  5. A vaguer wording does not just move the sources, it moves the question. “Où étudier à Paris ?” returned libraries, coworking spaces and city guides alongside universities.

What we measured

IntentsChoosing a private university in Madrid, choosing a university in Paris
Wordings4 per market, in the market’s own language
MarketsSpain and Spanish, France and French, each set explicitly
SurfacesGoogle AI Overview, Google AI Mode and Gemini, pooled per wording
Runs2 per wording, recorded per wording rather than assumed
CacheBypassed. Every run is a fresh upstream call
Date10 August 2026

The first wording in each market is not new. ¿cuáles son las mejores universidades privadas de Madrid? and Quelles sont les meilleures universités de Paris ? were already in our register from earlier work, so the baseline is a string we had published rather than one written for this study.

Two rows carry a note. The French second wording first came back reporting two runs while every surface had failed one of them, which makes each source stable over a single successful call. That is the same trap as a one-run reading and harder to spot, because the run count looks right. It was re-measured clean and both rows are in the register.


The eight cells

MarketWordingStable shareCitedHeld
Spain«¿qué universidad privada de Madrid tiene mejor reputación?»56%169
Spain«mejores universidades privadas de Madrid»47%199
Spain«¿dónde estudio en Madrid en una universidad privada?»43%219
Spain«¿cuáles son las mejores universidades privadas de Madrid?»39%187
France«Quelles sont les meilleures universités de Paris ?»53%1910
France«Quelle université parisienne a la meilleure réputation ?»44%167
France«meilleures universités parisiennes»32%258
France«Où étudier à Paris ?»30%3310
cited in every rundistinct sources citedES · mejor reputación9 → 16ES · mejores privadas9 → 19ES · dónde estudio9 → 21ES · cuáles son7 → 18FR · meilleures de Paris10 → 19FR · meilleure réputation7 → 16FR · parisiennes8 → 25FR · où étudier10 → 33
Each row is one wording of the same intent. The left dot is what held across both runs, the right dot everything cited at all. Google AI Overview, AI Mode and Gemini pooled, cache bypassed, 10 August 2026.
Stable sources against all cited sources for four wordings of a university question in Spain and France
cited in every rundistinct sources cited
ES · mejor reputación916
ES · mejores privadas919
ES · dónde estudio921
ES · cuáles son718
FR · meilleures de Paris1019
FR · meilleure réputation716
FR · parisiennes825
FR · où étudier1033

The percentages are not the finding here. Spain sits in a 17 point band and France in a 23 point band, which is unremarkable next to the 57 points the American hotel question moved. The finding is what happens when you intersect the four sets instead of averaging them.


Spain kept a core. Travel did not.

Same country, same method, same dayDomains stable across all four wordings
Hotel booking0 of 61 seen
Choosing a private university4 of 15 seen

That single comparison is why this study exists. This morning’s result showed four Spanish hotel wordings returning a respectable looking 42%, 41%, 48% and 30% while sharing not one reliably cited domain, and the tempting conclusion was that a stable core is always an artefact of the sentence. It is not. Change the category and the same instrument finds a core that four different sentences agree on.

The pairwise overlaps say it too. In Spanish education the four wordings share between 0.29 and 0.78 of their stable sets, measured as intersection over union. In Spanish travel the same figure ran from 0.06 to 0.32. The two closest education wordings agree about three quarters of what they cite; the two closest travel wordings agreed about a third.

So “does rewording break your visibility” has no general answer, and a vendor quoting one is averaging over categories that behave nothing alike. It is a per-category, per-market measurement, and it is cheap enough that nobody has an excuse for guessing.


The core exists, and universities are not in it

Here is the part an education marketing team should read twice.

The four Spanish domains that held across every wording are graddus.com, uniscopio.com, smartresidences.es and maclapiovera.es. Two are comparison and ranking sites, one is a student housing operator and one is a local guide. None is a university. In France, the single survivor is relocation-in-paris.fr, a relocation service.

Universities do appear, and they appear conditionally. ucjc.edu, udit.es and cunef.edu held for the wording that asks where to study, and for none of the other three. pantheonsorbonne.fr held for two Paris wordings and u-paris.fr for two others, never all four.

We have seen a related shape from the other side. Asking which agency is best across seven industries returned 133 distinct domains of which only ten were cited in more than one sector, and a single domain was the answer in five of the seven. Durable positions in these answer sets are rare and they are held by few domains, whether the question is about agencies or about universities.

That pattern lines up with what we found measuring universities named in AI answers but not cited as sources. The institution can be the answer while somebody else is the citation. This study adds the mechanism: the citation slot is durable and the institution does not own it. An intermediary that ranks or compares institutions holds a position that survives rewording, and the institutions rotate through the remaining slots depending on the exact sentence. We have since counted the same thing across eleven commercial categories, and it holds hard enough to be quantified: fewer than three in ten durable citations belong to a company that sells the product being compared, and in four categories the number is zero.

For an institution the practical reading is uncomfortable but clear. Competing for the durable slot means publishing the comparable thing, the ranked list, the structured table, the guide with the numbers in it. Competing for the rotating slots means covering more than one wording, because the wording is what selects you.


A vaguer wording changes the question, not only the answer

The lowest French cell, “Où étudier à Paris ?”, is the most instructive of the eight. It returned 33 distinct sources, the largest set in the study, and its stable set overlaps the other three wordings by 0.18, 0.13 and 0.06.

Reading the sources explains why. Alongside diplomeo.com and icd-ecoles.com it returned a childcare service, a rental guarantor, a student housing marketplace, the Paris city portal, the municipal library site and a coworking space. In French, “où étudier” can mean which institution to enrol in or where to physically sit and study, and the engine answered both.

That is not noise. A wording that is ambiguous to a human is ambiguous to the engine, and the answer set splits accordingly. If a prompt set contains one of these, the resulting number describes a blend of two questions and no amount of repetition will separate them.


We measured what a second reading is worth, and then took five

Every study we publish carries a caveat saying two runs is an easier bar than three, and puts that difference at about ten points. Nobody had measured what simply asking again is worth, so we re-ran three of the four Spanish wordings in a second session the same day, same three surfaces, same two runs, nothing changed but the clock. They moved by 18, 4 and 7 points, and their stable sets still overlapped 0.67, 0.67 and 0.70.

Then we took the first wording and read it five times.

ReadingSources citedHeld every runPublished share
118739%
214857%
311764%
413862%
59889%

One question, one day, one method: the share runs from 39% to 89%, and the number of durable sources never leaves seven to eight. The metric moves 50 points. The thing it claims to measure moves by one source.

All of the movement is in the denominator. What changes between readings is how many one-off sources happen to turn up, from nine to eighteen, and dividing a near-constant numerator by that produces a percentage that looks like a measurement and behaves like a coin.

And the core is not just stable, it is the same core. Four domains were cited in every run of all five readings, graddus.com, smartresidences.es, maclapiovera.es and uniscopio.com, and they are the same four that survived all four wordings above. Survives rewording, survives re-measurement, same four names.

One thing we cannot settle, and it is a caveat against our own instrument. The totals fall across the sequence, 18, 14, 11, 13, 9, and the readings were taken in order, so we cannot separate ordinary variation from repeated measurement shrinking the set. A different wording measured in the same window returned 19, 20 and 17 across three readings, which does not show the same decline, but three readings do not settle it either. It is in the register as an open question rather than a footnote.

The practical instruction is short and it costs a vendor nothing to follow. Report the domains, not the percentage. A named list of four sources is reproducible across rewordings and re-readings. The percentage on top of it is not.


What this study does not show

  • Two intents, two markets, eight comparable cells. Every percentage has a denominator between 16 and 33.
  • Two runs, not three. “Cited in every run” is an easier bar over two runs, so these are not comparable to three-run measurements, including parts of our agency study, which put that effect at about ten points. The section above shows a second session is worth about the same, so the two sources of noise are comparable and neither is separable from the other in a single reading.
  • One vertical pair, not a survey of verticals. We now have travel and education. Two categories differing is enough to refute a universal claim and not enough to predict a third.
  • City-scoped questions. Both intents name a city, which narrows the candidate set and may itself favour a durable core. A national question might behave differently and we did not test one.
  • Three surfaces, not four. ChatGPT is excluded because the route we use returns no source list for it, and reporting that as a zero would be reporting our instrument.
  • We did not check search volume for the four wordings. This says the choice of wording matters, not which wording to choose. Volume would not settle it either: we measured separately that the tools with the most search demand are not the ones AI names, so picking the highest-volume phrasing is not the same as picking the one that gets you cited.
  • Nothing here is causal. We changed the words and recorded what came back.

Common Questions About This Study

Does rewording a question always change which sources AI cites?

No, and that is the finding. In Spain, four wordings of a hotel booking question shared zero reliably cited domains, while four wordings of a private university question shared four. Same country, same three surfaces, same method, same day. Whether rewording breaks your visibility is a property of the category, and it has to be measured per category rather than assumed.

Which sites get cited for university questions in AI answers?

In our Madrid measurement the domains cited in every run of all four wordings were two comparison sites, a student housing operator and a local guide. Universities were cited, but each one only for particular wordings: three institutional domains held for the wording asking where to study and for none of the other three.

Why do universities not appear in the stable set?

We can say what the data shows and not why. The durable citations went to sites that rank, compare or contextualise institutions rather than to the institutions themselves, which is consistent with our earlier finding that universities are frequently named in AI answers while somebody else is cited as the source. This study adds that the citation slot itself is stable across rewordings, and institutions rotate through the remaining ones.

How many prompt wordings should an education marketing team track?

More than one, and the number depends on the market. In Spanish the four wordings overlapped between 0.29 and 0.78, so two well chosen ones cover a lot. In French the overlaps ran from 0.06 to 0.50, so the same two would miss most of the picture.

Which AI surfaces were measured?

Google AI Overview, Google AI Mode and Gemini, pooled per wording, with the cache bypassed so every run is a fresh call. ChatGPT is deliberately absent: the route used here returns no source list for it, and a zero produced by an instrument is indistinguishable from a real zero.

Where can I see the raw rows?

In our measurement register, which carries every row of this study with its prompt exactly as it was sent, its intent, its market, its surfaces, its run count and its date, including the row excluded for completing only one run and the reason.


Where this leaves you

Two studies, run the same way a few hours apart, disagree about the thing everyone in this category asserts. One category loses its entire stable core when you rephrase the question. The other keeps a core of four, and hands it to intermediaries rather than to the institutions being asked about.

The instruction that survives both is the same one, and it is not a slogan. Measure your own category, with more than one wording, and look at which domains hold rather than at the percentage. The percentage was similar in both Spanish blocks. The thing underneath it was not remotely the same, and only one of those two facts tells you what to publish next.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)