Skip to content
AI VisibilityGEO
EN

In Spain, Google's AI Overview Answers Every Software Question We Asked and No Product Question

AI Overview returned nothing for 11 buying questions in Spain and Germany. AI Mode answered all 36 cells in three markets: the surface, not the market.

· Updated · 25 min read

If you sell a physical product outside the United States, the Google surface that every AI visibility tool measures may not exist for your category. We asked 84 buying questions in four markets. In the United States AI Overview answered all twelve. In Spain it answered all six software questions and none of the six product ones. Germany fell between the two: four of its six product questions came back empty, and one of its six software questions. Then we asked the same cells on AI Mode in all three markets, and it answered 36 of 36.

We publish this as a measurement of the market and not of any one company, and the whole design was written down before the first call, including the condition that would have killed it.

The short version

  1. In Spain, AI Overview did not appear for a single product question. Running shoes, noise cancelling headphones, a mattress, an espresso machine, a cabin backpack, a robot vacuum: six questions, six empty answers.
  2. In the same market it appeared for every software question. Invoicing, project management, accounting, electronic signature, email marketing, backup: six questions, six answers.
  3. In the United States it appeared for all twelve, the same six product questions in English and the same six software ones.
  4. Germany and France are neither of the other two. 4 of Germany’s 6 product questions came back empty and 1 of its 6 software questions; France declined 3 of 6 product questions and 0 of 6 software ones. Four markets, four different answers.
  5. It is the surface, not the market. The same cells asked on AI Mode the same day answered 36 of 36 across all 3 markets it was asked in, product and software alike. European product brands are not invisible to Google’s AI. They are invisible to the one surface every tool measures.
  6. The most extreme number in the study survived a re-read. Asked again later the same day, 6 of 6 product cells came back empty again, this time across 18 completed runs, while 5 of the 6 software cells answered in the same window. The software side flickers and the product side does not.
  7. The question that started this has been read 11 times. Two readings of 3 runs each in English, all answered, and 9 readings in German and Spanish, 6 of them empty. The 3 that answered are the ones that matter: a German AI Overview reading with a string that had returned nothing twice, and both AI Mode readings.
  8. This is about whether the surface shows up at all, not about who gets cited in it. If it does not show up, there is nothing to be cited in.

What was measured

84 cells, one surface per call, on 23 August 2026. Six physical-product questions and six software or service questions in each of three markets on AI Overview, Spain in Spanish, Germany in German and the United States in English, the same twelve intents translated, plus the same twelve in every market again on AI Mode. Each call asks the same surface two or three times and reports how many of those runs came back with an AI Overview. The 12 Spanish AI Overview cells were read a second time later the same day, so the register carries 96 readings of 84 cells.

QuestionsAI Overview never appeared
United States, physical product60
United States, software60
France, physical product63
France, software60
Germany, physical product64
Germany, software61
Spain, physical product66
Spain, software60

The register carries every cell with its prompt, the number of runs asked and the number answered.

Correction, 23 August 2026. This study published earlier today with two American software questions instead of the six the design named, because the other four timed out while their Spanish counterparts answered. They were read later the same day and all four returned an AI Overview, so the American software arm is now six of six. The finding did not change. What changed is that the half a reader could most easily dismiss is no longer the thin one. One of the four deviates and says so in the register: the electronic-signature cell was read as “software” rather than “service” after the designed string timed out three times, and with two runs instead of three.

The hypothesis we froze, and why it is retired

The design came from an accident. In an earlier study, one intent kept coming back with nothing: running shoes for beginners, in German and in Spanish, four readings, not one AI Overview, while the software questions around it answered normally.

The obvious explanation is that Google declines the product recommendation question because its results page already has a Shopping unit doing that job. That is a claim about a class of question, so the design said: ask six product questions in the United States, and if a fifth or fewer of them decline, the hypothesis is retired.

None of them declined. Six product questions in English, six AI Overviews. By the band written before the first call, the class of question is not the answer, and the hypothesis as stated is retired.

What replaced it is more useful, because it is the same class behaving differently in two places: the product questions decline in Spain and answer in the United States, and the software questions answer in both. The class only matters inside a market.

The third market, and the reason it was worth twelve calls

With two markets a split like this cannot name its own outlier. The United States answers everything and Spain declines its product half, and that sentence reads as a fact about Spain only if you already assume the United States is the normal one. It is exactly as consistent with the United States being the unusual one, and nothing in a two-market table tells you which.

We had run into the same shape in another study the day before, where the third market moved the outlier from one side to the other and cost a headline that had gone out that morning. So the German arm was designed and written down before the first call, with the bands in it: 0 or 1 of six German product declines would make the finding Spanish, 5 or 6 would make it European and the headline wrong, and anything in between would be published as partial.

Germany came back at 4 of 6. By the band written in advance that is partial, and it is neither of the other two results. So the finding is not that Spain is strange. It is a gradient, and the honest way to write it is with three numbers instead of one.

The control is what makes that gradient mean anything at all. If Germany had declined its software questions too, this arm would not be measuring the class of question, it would be measuring that Germany declines everything. It declined 1 of 6, inside the band the design named for a working control, so the split between product and software exists in Germany as well and is simply weaker there.

And one intent has now been seen on both sides. Running shoes for beginners, in German, returned nothing in two readings on two different days and returned an AI Overview today, from the same string. In Spain the same intent has returned nothing in 4 of 4 readings. That is the gradient at the level of a single question: in Spain the absence repeats, in Germany it does not, and a German brand that measures once can get either answer.

The fourth market, and what it settles

With three markets the result reads as a gradient and nothing inside it can say whether that gradient is European or whether Spain is simply the outlier. Only a fourth market can, and the prediction and its bands were written down before the first call: 4 or more product declines in France would make the class split European, 1 or 0 would make it a Spanish and German finding, and 2 or 3 would be published as partial.

France declined 3 of its 6 product questions and 0 of its 6 software ones.

So it is partial, and the useful shape is the order rather than any single number. Product questions with no AI Overview at all: 0 of 6 in the United States, 3 of 6 in France, 4 of 6 in Germany, 6 of 6 in Spain. Software questions with no AI Overview: 0, 0, 1 and 0. The class split shows up in all three European markets and in none of the American ones, and what changes between them is how hard it bites.

The control is what makes that readable. If France had declined its software questions too, this arm would be measuring that France declines everything rather than that the class of question matters. It answered all six, which is the band the design named.

One honest weakness, because short reads bias in a single direction here. Two of the twelve French cells stopped early, and a cell that returns nothing across two runs has had fewer chances to appear than one that returns nothing across three. The mattress question is the one that carries it: it declined across 2 runs rather than 3. The other two declines ran the full three.

The same questions on the other surface

This piece told a Spanish product brand to move its work to the surfaces that do answer. That sentence was leaning on cells from a different study, which makes it the weakest-supported claim here and the only one that tells a buyer what to do. So it was worth twelve calls, and the design was written down before the first one: if the buyer really has somewhere to go, AI Mode declines 0 or 1 of the six Spanish product questions; if it declines 4 or more, the advice is wrong and there is no Google surface answering product questions in Spain, which would be a bigger and much worse finding.

AI Mode answered 12 of 12. Six product questions, six software questions, the same strings, the same market, the same day. Not one empty answer.

Then the same twelve in Germany, because a claim about surfaces measured in one market is a claim about that market until somebody checks. Germany was the one worth twelve more calls: its AI Overview baseline is 7 of 12, neither floor nor ceiling, so the comparison could have moved either way. In the United States it could only have tied, which is a test that passes by tie and tells you nothing.

It did not move. AI Mode answered 12 of 12 there too.

Same 12 cells per market, same dayAI OverviewAI Mode
Germany7 of 1212 of 12
Spain6 of 1212 of 12

That is 36 of 36 for AI Mode. AI Overview left 11 of those cells empty and AI Mode answered all 11, every one of them, in two markets and two languages. Not one of those eleven questions is unanswerable by Google’s AI. They are unanswered by the surface everybody measures.

And the third market closes the surface gap. The twelve American cells were asked on AI Mode the same way, with the prediction and its threshold written down before the first call: 11 or 12 answering would leave “AI Mode answers where AI Overview does not” a property of the surface, and 3 or more declining would have cut the claim back to Europe. It answered all 12. The design also asked whether the class of question shows up on this surface at all, and it does not: 0 of its 6 product cells and 0 of its 6 software cells came back empty, a difference of nought where two or more would have meant a class effect on AI Mode as well. Those cells cite a median of 11 distinct domains each, which is the first breadth figure this study carries on that side, recorded without a threshold so a later arm has something to compare against.

The intent that started all of this makes the point on its own. Running shoes for beginners, in Spanish, has returned nothing from AI Overview in 4 of 4 readings. On AI Mode it answered all three runs and cited sixteen domains: Decathlon, Runnea, RunRepeat, Running Point, a specialist review site, Reddit, YouTube and TikTok among them.

So a Spanish product brand is not invisible to Google’s AI. It is invisible to the surface that every AI visibility dashboard reports on, while the sibling surface answers the same question with a page full of retailers and reviewers it could plausibly be in.

The control matters here as much as the result. If AI Mode had answered the product questions and declined the software ones, this arm would be measuring something other than the surface. It declined neither, 0 of 6 product and 0 of 6 software in Spain, and nothing at all in Germany, which is the cleanest version of the control the design could have got.

We read the most extreme number again

The strongest claim here is that six Spanish product questions returned no AI Overview at all. That is exactly the kind of number worth doubting, and the rule this register works to is that the thing to re-measure is your best number: it is the cheapest to check and the one that changes the conclusion most if it falls.

So the same six were asked again later the same day, and so were the six software questions, because a second empty reading on its own cannot be told apart from AI Overview simply being down.

6 of 6 product cells came back empty again, this time across 18 completed runs. Not one AI Overview in any of them. And the surface was up while we asked: 5 of the 6 software cells answered in the same window.

The sixth software cell is the interesting one. Project management answered both runs of the first reading and none of the second. So the software side flickers and the product side does not, which is a sharper way to put this study’s finding than either number alone: across two readings the product questions never produced an AI Overview, while the software questions that do produce one do not produce it every time.

That also sets the honest way to read the table above. A software cell marked as answering is a cell that answered when we asked, not a cell that answers always. The product cells are the ones where two readings agree.

And the assistants cite sources for the questions AI Overview refuses

The obvious next question from a Spanish product brand is: fine, but what about the assistants people actually use? This piece could not answer that and said so, which made it twelve calls.

The statistic has to change here, and saying why is half the work. “Did the surface appear” only means something where the surface can fail to appear. An assistant always answers. What can be missing is not the answer, it is the sources. So this arm counts distinct domains cited rather than blocks returned, and it sits in its own section and its own part of the register for exactly that reason.

The same twelve Spanish questions, asked of Gemini, the same day. All 12 of them cite sources. The 6 product questions AI Overview refuses come back with between 5 and 11 distinct domains each, and the 6 software questions do too. Not one comes back empty: 0 of 6.

Which settles the thing the study could not settle on its own. The six questions AI Overview leaves empty are not unanswerable, and they are not unanswerable with sources either. They come back with Spanish retailers, consumer titles and specialist reviewers inside them: Decathlon, El Corte Ingles, Jysk, Comparamaletas, RunRepeat, La Vanguardia, ABC, El Independiente. One Google surface declines to do what its own sibling and its own assistant both do.

One unit warning travels with that number. These are distinct domains and never citations. Gemini anchors each citation to a passage inside a page, so several citations resolve to the same page and the two units differ by roughly a factor of three. A citation count from Gemini set against a page count from another surface is comparing two different things.

And France carries the comparison Spain cannot. Spain refuses all six product questions, so inside that market there is nothing to compare them against. France refuses three and answers three, which makes it the one place in this study where the market, the language and the day can be held still while only AI Overview’s decision changes. Asked of Gemini, the 3 refused cells cite sources in 3 of 3 and the 3 accepted product cells cite sources in 3 of 3, at a median of 4 distinct domains each against 4.

That rules out the cheapest explanation of the whole study. If AI Overview were declining the questions with less material behind them, the refused cells would come back thinner than the accepted ones. They come back the same. Whatever decides it is not how much citable material exists about the question.

Across the 12 French cells Gemini cites a median of 6 distinct domains against 7 in Spain, and on product questions alone the French range is 4 to 8 against 5 to 11.

And ChatGPT is absent on purpose. Through this route it returns no source list at all, so a zero from it would be a zero of the instrument rather than of the world, and this register does not publish those.

The candidate mechanism, measured and retired

Every version of this study has carried the same caveat: the Shopping unit is the explanation the pattern suggests, and this design does not observe it. It has now been observed, through a different route from everything above. The same twelve Spanish questions were read as an ordinary organic result page, and what was recorded is whether Google put a product block on it.

The retirement threshold was written down before the first call: 2 or fewer of the six product cells carrying a block retires the candidate. 2 of 6 carry one. Noise cancelling headphones has a popular products block and a product-sites block; the cabin backpack has a popular products block; running shoes, the mattress, the espresso machine and the robot vacuum have none.

AI Overview declined all six. So a product block cannot be why it stayed away from the other four, and the candidate is retired rather than confirmed. That is worth saying plainly: this study now has no candidate mechanism at all, and the honest position is that we measure where the surface is absent and not why.

A zero from this sweep would have two readings, no block on the page or no block the instrument can see, so a positive control ran first and returned a popular products block twice and a product-sites block on a query that certainly carries one. 3 of the software cells returned a provider error and are absent from every count rather than counted as zero.

And something fell out of the same response that this arm was not designed for. The organic page carries an AI Overview item in 0 of the 6 product cells and 3 of the 3 readable software cells. That is the finding at the top of this piece, reproduced by a second instrument that knows nothing about the first.

The control that rules out our own instrument

A finding like this is worth exactly as much as the checks against its boring explanations, so here they are.

It is not the shape of the call. The earlier readings asked two or three surfaces at once. The English run that answered three times out of three asked one surface. So the Spanish prompt was asked again with that same single-surface shape, and it returned nothing. Same instrument, same call, opposite result.

It is not a timeout. Asked through a second, separate route, the Spanish prompt returns a successful response with an empty body. That is the surface answering the request and producing no AI Overview, which is a different thing from a request that never came back. Cells that genuinely timed out are absent from the register rather than recorded as zeros, because an instrument zero and a real zero look identical in a table and mean opposite things.

It is not one unlucky question. Six different product categories in Spain, six empty answers, and the one we have asked most often has 6 empty readings in German and Spanish against 3 answered runs in English.

What this costs a buyer

If you sell physical products outside the United States, the surface everyone measures may not be in your category at all. Not that you rank badly in it. That it does not appear.

That changes three things about how you would spend.

  • A dashboard that reports “your AI Overview visibility” for a Spanish product catalogue is reporting on something that did not appear. Ask any vendor how their number behaves when the surface returns nothing: a share of zero and no data are different states and a percentage hides the difference.
  • The work moves to the surface that does answer, and that is now measured rather than assumed. The same cells on AI Mode answered 36 of 36 across the three markets. That is where a product brand in Spain still has something to win, and the answers there cite retailers, consumer associations, newspapers, reviewers and video.
  • And you cannot check this by looking at English. The same question, translated, gives the opposite answer. Any audit that samples one market and generalises will get this backwards.

This is the kind of thing our own measurement is built to catch, and if you want the same question asked for your categories and markets, that is what our platform does.

What we would have concluded from English alone

This study started as the opposite claim, and it is worth being explicit about how close it came to being published that way.

The first four readings all came from European markets, so the pattern in front of us was “AI Overview declines product questions”. The design we froze went looking for that in the United States, where it is easy and cheap to test, and the United States said no six times out of six. Had we run the English arm first and stopped there, the conclusion would have been the mirror image: product questions are fine, and the earlier European blanks were noise.

Both readings are wrong on their own, and the thing that separates them is a single translated question asked in two places. That is the smallest useful unit of an AI visibility audit and almost nobody runs it, because a dashboard is built around one market at a time.

How to check your own categories in five minutes

You do not need a platform to find out whether this affects you, and the procedure is short enough to run by hand.

  1. Pick five buying questions your customers actually type, in the language they type them in, not translated from your English keyword list.
  2. Ask each one two or three times in the market you sell in. Once is a coin toss: an AI Overview that appears sometimes and an AI Overview that never appears look identical in a single look.
  3. Record whether the block appeared at all, before recording anything about who is in it. That binary is the finding here, and it is the one most tools skip straight past on their way to a percentage.
  4. Then do the same five questions in English. If the English versions answer and yours do not, you have the split this study measured, and your AI visibility work belongs on another surface.

If the block never appears, no amount of content work on that page changes anything about it, which is the most useful thing a measurement can tell you before you spend.

What this does not show

  • 4 markets and four languages. The United States in English, Spain in Spanish, Germany in German and France in French, twelve intents each. Nothing here says what Italy or the Netherlands would do.
  • 2 surfaces, and only one of them in every market. AI Overview in all three markets, AI Mode in Spain and Germany. Gemini has its own section and its own statistic, on the twelve Spanish cells only. Nothing here says what AI Mode does in the United States, and ChatGPT is absent because this route returns no source list for it.
  • One day, 2 readings. 23 August 2026, and the Spanish cells asked twice within it. AI Overview coverage is a product decision and product decisions change.
  • Two to three runs per cell. Enough to separate “never appeared” from “appeared”, which is the whole statistic here, and not enough to call anything a rate.
  • The Shopping unit was the candidate and it is retired, not replaced. It has now been observed: 2 of the 6 product cells carry a product block and AI Overview declined all six. We still report that the surface is absent, not why. Retiring a candidate is not finding a mechanism.
  • Nothing here is causal. We measured what appeared, not what makes it appear.

Frequently asked questions

Does this happen in Germany too?

Partly, and the word is doing the work. 4 of the 6 German product questions returned no AI Overview and two did, against six of six empty in Spain. 1 of the 6 German software questions also came back empty, against none in Spain. So the same split exists in Germany and is weaker there, which is why this piece reports a gradient across three markets rather than a rule about one.

Does AI Overview really never appear for product searches in Spain?

In this measurement, six product questions in Spanish returned no AI Overview across every run we asked, and the same six came back empty on a second reading later the same day. That is a strong and narrow claim: six categories, one day, one market. It is not a statement that the surface is switched off for every product query in Spain, and the honest way to use it is to ask your own categories rather than to assume.

Could this be an artefact of the measuring tool?

It is the first thing we checked. The same Spanish prompt was asked with the identical single-surface call that returned three answers out of three in English, and it returned nothing. A second, independent route returns a successful response with an empty body, which is the surface declining rather than the request failing. Cells that genuinely failed are absent from our register instead of being recorded as zeros.

If AI Overview does not appear, where should a product brand work instead?

On AI Mode, which answered 12 of 12 of the same Spanish cells on the same day, product questions and software questions alike. Those answers cite ordinary pages: retailers, consumer associations, newspapers, specialist reviewers, forums and video. The practical move is to find which surfaces answer your questions in your market before deciding where to spend, which is exactly the measurement this study is made of.

Why does the same question behave differently in English?

We do not know, and saying so is the point. The Shopping unit is the obvious candidate, because a results page that already answers “which product should I buy” with a carousel has less reason to answer it again in prose, and Shopping coverage differs by country. This study observes the absence and not the cause, so we are naming that as a candidate rather than a mechanism.

How many times did you ask?

One to three runs per cell across 84 cells in 4 markets and 2 surfaces. The 12 Spanish AI Overview cells were read a second time later the same day, so the register carries 96 readings of 84 cells. The intent that generated the hypothesis carries 4 more readings in German and Spanish from two other studies in this register, and 2 English readings of 3 answered runs each here. The statistic is binary per cell, whether the surface ever appeared, so the number of runs matters only in that it must be more than one. Every prompt, market, run count and answer count is in our measurement register.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)