Skip to content
AI VisibilityResearch
EN

We Found a GEO Tool's System Prompt in Our Search Console

For eight days in June, 163 Google queries hitting our site carried an identical instruction preamble. A monitoring tool was leaking its own method.

· Updated · 18 min read

We connected Google Search Console to our own reporting this week and read three months of query data for the first time.

One hundred and sixty-three of the queries that showed our site began with the same sentence:

analyze the following query for geo visibility data. mandatory rules: 1) use live web search before producing the final answer. 2) use at least 3 distinct sources when available. 3) include source urls in your answer (not placeholders). 4) do not return only instructions/template; return extracted findings. provide: (1) ordered mentions, brands/pages/companies in the sequence they appear, (2) key entities, important topics, concepts, and named entities, (3) response structure, headings and sections used. query: …

That is not a person searching. That is an AI visibility monitoring tool, running its customers’ questions through Google with its entire instruction template pasted into the search box.

Disclosure: EchoWi sells AI visibility monitoring, so this is a competitor’s method arriving in our own data. We cannot identify which tool it is and we are not guessing. Figures are from our Search Console property for 3 May to 3 August 2026.


The short version

  1. 163 queries, all carrying the same instruction preamble, and 17% of the quarter’s impressions.
  2. Every one of them fell in an eight-day window, 25 June to 2 July 2026, with 42% of them landing on 30 June alone and nothing before or after.
  3. That was 31% of all our non-brand impressions for the quarter.
  4. The embedded questions belong to somebody’s clients, across at least a dozen unrelated industries.
  5. The wider number is worse: across three months, not one non-brand query produced a single click.
  6. You can find these without reading a query at all. Across 6 query families measured over 12 days, the two templates we had already identified put 97% and 85% of their impressions into their three biggest days, against 48% to 58% for ordinary search behaviour.

What the numbers are

QueriesImpressionsClicks
Brand and misspellings5%44%100%
Non-brand95%56%0%
Of which, the tool’s preamble39%17%0%

Google Search Console, echowi.ai, 3 May to 3 August 2026. Shares of the quarter, with the totals withheld: we sell visibility measurement and our own absolute numbers are nobody’s business but ours. Every finding below survives the redaction, which is the test of whether it was a finding.

Read the clicks column twice. Every single click this site received in three months came from someone typing our name or misspelling it. Ninety-five per cent of our queries, and more than half of our impressions, produced no click at all.

That includes 35 queries where we ranked in the top position. Ranking first and getting nothing is not a CTR problem you fix with a better title.


The eight days

The preamble queries did not trickle in. They arrived as a burst and stopped.

DateImpressions
25 junio15%26 junio0%27 junio3%28 junio5%29 junio14%30 junio42%1 julio14%2 julio6%
Los ocho dias, como reparto de la ventana. Empieza el 25 de junio, hace pico el 30 y para el 2 de julio. Los denominadores se retiran a proposito: lo que importa es la forma, no nuestro tamano.
Share of the tool's impressions by day across the eight-day window
value
25 junio15%
26 junio0%
27 junio2.9%
28 junio5.2%
29 junio14.4%
30 junio42.3%
1 julio14.1%
2 julio6.1%

Something started on 25 June, ran for eight days, and stopped. A trial that ended, a misconfiguration that got fixed, or a campaign that finished. We have no way to tell which, and the tool has not been back since.


What was inside the template

The instruction block is identical every time. What changes is the question pasted after query:, and those questions are not about us or about our category at all.

Across the 163 we counted questions about branded land developments and plotted real estate in India, enterprise email and collaboration suites, AI advertising platforms, business messaging in telecoms, restorative dentistry, integrative wellness clinics, commercial LED lighting, agricultural shade netting, screen enclosures for hurricane weather, secondhand fashion marketplaces, artificial surf machines, corporate t-shirt sourcing and online chess tuition for children.

One vertical dominates. More than fifty of the queries concern a single Indian real estate developer by name, asking about specific projects, refund policies, price comparisons against named competitors and whether particular plots are a good investment.

We are not naming any of them. Those brands did not choose to appear in our Search Console, and publishing them would expose which vendor they use and which questions that vendor thinks decide their sales. We found the hotel chain in an earlier study the same way and did not name it either, for the same reason.

What is fair to report is the shape, because the shape is the finding: one tool, one template, a client roster spanning at least a dozen unrelated industries, and a habit of putting the whole thing into a search box.


Why the template ended up in Google

The tool is doing something reasonable in an unreasonable way.

To measure whether a brand appears in AI answers, you have to ask the question and read the response. Some tools query the assistants directly. Others go through a search surface, because AI answers are increasingly assembled from live web results and because a search engine is cheaper and more stable than a chat interface.

What appears to have happened here is that the whole prompt was passed as the search string rather than only the customer’s question. The instruction block is meant for a language model. Google received it, treated it as a very long query, matched it against pages, and dutifully logged an impression for every site it showed.

That is why we appear. Our articles are about GEO visibility data, ordered brand mentions and source extraction, which is exactly what the first two hundred characters of the preamble talk about. We were matching the instructions, not the question.


What this costs the tool, and what it costs you

For the tool, three things leak at once. Its method, since the mandatory rules describe how it extracts mentions and entities. Its client list, by industry and in several cases by brand name. And its measurement, because a query padded with three hundred characters of instructions does not return the same results as the customer’s question on its own. Whatever those runs measured, it was not what the client asked.

For everyone else, the cost is quieter and more general. If you publish in a category that AI visibility tools monitor, some fraction of your Search Console impressions is robots. Ours was 31% of non-brand impressions for a quarter, produced by one tool over eight days.

That matters because impressions are the metric people use when clicks are too few to trust, which is exactly the situation most sites are in right now. A rising impression count that is really a competitor’s crawler is worse than no data.

How to check your own, in about ten minutes:

StepWhat to look for
1Sort Search Console queries by length. Anything over 200 characters is not a person
2Look for imperative verbs: “analyze”, “evaluate”, “compare the top”, “provide”
3Look for scaffolding: “mandatory rules”, “do not”, “context:”, “question:”, “be neutral”
4Check the dates. Robots arrive in bursts, humans do not
5Check the clicks. A query with impressions, a good position and zero clicks over months is a machine

We found more than the one template, incidentally. Other queries in the same period carried fragments like “do not update memories”, “be neutral: do not assume it relates to seo/marketing unless confirmed”, and “context: location: united states (not for language). do not include location references in your response.” Different scaffolding, same phenomenon, smaller volume.


The finding under the finding

Strip the robots out and the picture does not improve, it clarifies.

Every non-brand query. Three months. Not one click.

Some of those queries we rank first for, and they are exactly the questions our articles were written to answer: how to measure brand visibility in AI search, which platforms compare a brand against competitors in AI answers, what to evaluate first when AI describes your product wrongly.

We are first, and nobody arrives.

Part of that is that some of them are machines. Part of it is that a question-shaped query in 2026 often gets answered above the results, by the same AI answer layer this whole site is about, and the person never scrolls. We have measured that layer from the outside for months. This is the first time we have seen the other side of it, in our own data, as an absence.

What we would not conclude from this. That ranking is worthless, or that this is proof of anything about the wider web. It is one small site in one young category over one quarter, and a site this small has no business generalising about search. What we can say is narrower and still useful: in this category, at this size, position and impressions have stopped predicting visits, and any reporting built on them is measuring something other than people.



A second template, and the filter that could not see it

Three days after this study went up, a fresh twenty-eight day window turned up a second one. It comes from a different tool, and its shape has nothing in common with the first.

The first template announces itself. It opens with instruction words, numbers its rules, and demands live web search and at least three distinct sources. It is unmistakable once you have seen it, and the detection rule we published was built out of exactly those words.

The second one carries no instructions at all. It is a pairwise comparison with a machine-readable footer:

between {tool} and {rival}, for us - geo - transactionnel,
who is the best? who wins?
<context> company url: https://{tool}.ai market: france </context>

Seven different rivals appear in the second slot, and the same handful of judgement axes rotate through it: who wins, who is more reliable, who has superior customer support, who offers better AI search innovation. It is a vendor benchmarking itself against its competitors, one pair at a time, and the searches leak the same way the first one did.

The two do not overlap by a single query. The filter we published catches eleven rows in this window. It catches none of the second template, because it was built from the vocabulary of the first. Applying both, the count is thirty-six queries out of the one hundred and fifty-four that run to twelve words or more, which is twenty-three per cent of the long tail rather than the seven per cent the original filter reported.

That correction matters more than the discovery. Anyone reading the first version of this study would have written the same filter, run it, seen a small number, and concluded the problem had mostly gone away. We did, for a day.

So the rule is not the filter, it is the shape. An agent query is long, it is grammatical, it appears once or twice and never again, and it carries something no human types: a rules block, an XML-ish context tag, a placeholder that was meant to be substituted, a market name in a language that is not the query’s. Grep for the vocabulary you have already seen and you will find exactly the templates you have already seen.


The other shape, the one that does not age

Everything above is about the shape of the query string, and it has a shelf life. Every marker in it is a word some vendor chose, so the morning they reword the template the filter goes quiet and the count reads like good news. So we went looking for a marker no vendor controls.

A query family has a time series. A tool run has a duty cycle: it starts, it runs for a few days, it stops. Demand does not do that.

Measured across twelve days on six query families from this property, each normalised to its own peak day so the shape is visible without any total. The statistic is the share of a family’s impressions landing in its three biggest days, which is a ratio inside the family and needs no denominator from us.

Query familyThree biggest days
Vendor battery about one company, persona-prefixed97%
Pairwise comparison template with a fixed tail85%
Entity questions about one company in bulk61%
Brand58%
Queries containing reviews53%
Queries containing pricing48%

The two families we had already caught by their wording are the two most concentrated, and nothing else is near them. One sits at almost nothing, carries essentially everything it will ever carry across three consecutive days, and drops back to almost nothing. The three controls ramp up and stay up, which is simply what a site looks like while it is publishing more pages.

This is not an artefact of family size. The largest family is the most concentrated and the smallest is second, so the ordering is not the small numbers being spiky.

The third row is the one worth keeping. It asks entity questions about a single company in bulk, dozens of them, and it reads exactly like a battery. It does not have the shape of one: 61% against a control maximum of 58% is not a separation, it is two numbers next to each other. So the duty cycle is sufficient and not necessary. A family that switches on and off inside a few days is a tool run; a family that plateaus can be either, and no amount of staring at the time series will tell you which.

And the battery has grown a shape the vocabulary rule does not describe. The largest family asks the same buyer questions about one vendor in two languages, and the Korean half puts one of three rotating role declarations in front of every question: chief marketing officer of a Fortune 500 company, SEO lead, brand manager. A person does not restate their job title before each search. Nothing in that is a rules block, an XML-ish tag or an unsubstituted placeholder, so the published filter walks straight past it, exactly as it walked past the pairwise template.

The reason to measure the calendar rather than the wording is that a vendor can change every word of a template overnight and cannot change the fact that a batch run has a beginning and an end. You can run this on your own property without reading a single query string: group by family, divide each day by that family’s own biggest day, and look for a rectangle.

Two things this does not do. It is six families on one property over one window, so the boundary between 85% and 61% is where our data happens to fall and not a threshold anyone should hard-code. And two of the three controls are commercial queries about competitors, which is what most of our non-brand coverage is, so they are not independent of each other.


Where it lands, which is not spread out at all

Measured again three months out, with both templates in the filter, the preamble is thirteen per cent of the site’s impressions. Spread evenly that would be a rounding error on any single page, and the next number is why it is not.

Ninety-one per cent of those impressions land on one article. Not distributed across the blog: one page absorbs almost all of it, and on that page the preamble is forty-nine per cent of everything the page receives, against zero clicks.

Both of those percentages were wrong when this section was first published, and the way they were wrong is the same species of error the rest of this study is about. They were computed from an export grouped by query and page, which counts an impression separately for every page a query surfaced. Non-brand queries only ever show one page, so they are unaffected; a navigational brand query gets sitelinks and fans out across eighteen. The denominator more than doubled, the numerator did not move at all, and the share came out at six per cent instead of thirteen.

Nothing about the finding changed, which is precisely why it survived. A distortion that moves the headline number and leaves the conclusion intact does not look like a mistake from any angle, and the only thing that catches it is deriving the same figure a second way.

That page is, by impressions, the biggest on the site. Read from a dashboard it looks like the clearest opportunity we have: high impressions, average position just off page two, obvious room to climb. Half of it is a robot, and the queries inside the wrapper are about land developers in India, digital banking in Turkey and surf ride installations, none of which we have written a word about. We match because the wrapper says “geo visibility data”, not because anyone wanted us.

We made that mistake in our own notes before catching it, and filed the page as a depth problem to be fixed by ranking better. It is not a depth problem. There is nothing under it to rank for.

So the practical rule is that a bot’s traffic is concentrated, not diffuse. A site-wide share of a few per cent can still be half of your most promising page, and the average position on that page is an average over hundreds of machine queries that no person will ever type. Segment by page before you plan anything from an impressions column.

What this does not show

  • We cannot identify the tool. No user agent, no referrer, no IP. Search Console gives you the query string and nothing else, so every attribution here is about the template, not about a company.
  • We cannot prove the pattern is common. One property, one quarter. If you check yours and find nothing, that is a real answer and we would like to hear it.
  • The 31% is specific to us. A site in a category nobody monitors would see none of this, and a bigger site would see the same absolute volume as a smaller share.
  • Search Console counts are approximate. Google samples and filters query data, so treat every figure here as its reported value, not as a census.
  • Zero clicks is not zero value. Impressions in an AI answer layer may still lead somewhere we cannot see, and our own research says attribution in this environment is largely unsolved.

Common Questions About This Study

What exactly did you find in Search Console?

163 queries between 25 June and 2 July 2026 that all began with the same instruction block, telling a language model to analyse a query for GEO visibility data, use live web search, cite at least three distinct sources, and return ordered brand mentions, key entities and response structure. The customer’s actual question was pasted after the word “query:”. Together they produced 17% of the quarter’s impressions and no clicks.

How do I know if this is happening to my site?

Sort your Search Console queries by length and read the longest ones. Human queries are rarely over 200 characters. Look for imperative verbs, numbered rules, and phrases like “do not” or “context:”. Then check whether they arrived in a burst on consecutive days, because that is what automated runs look like and human demand does not.

Does this mean my impression numbers are wrong?

It means they may include machines, and you cannot tell without reading the queries. For us it was 31% of non-brand impressions for a quarter, from a single tool over eight days. If you report impressions as a growth metric to anyone, it is worth knowing what fraction of them has a pulse.

Why won’t you name the tool or the brands?

We cannot identify the tool: Search Console shows the query string and nothing that identifies who sent it, so naming one would be a guess presented as a finding. The brands are a different question. They did not choose to appear in our data, and publishing them would expose which vendor they use and what that vendor thinks decides their sales, neither of which is ours to give away.

Is querying Google a legitimate way to measure AI visibility?

It can be, since AI answers are increasingly built from live web results, and a search surface is cheaper and steadier than a chat interface. The problem here is not the approach, it is that the instruction block appears to have been sent as the search string. A query padded with three hundred characters of instructions does not return what the customer’s question returns, so those runs measured something other than what the client asked.

What should a monitoring tool do instead?

Send the customer’s question and nothing else to the search surface, and keep the instructions where they belong, in the model call that reads the results. It is worth checking: if your scaffolding is going into the query string, it is being logged by every site you touch, along with enough of your client list to reconstruct it.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)