Skip to content
AI VisibilityAgencies

We Asked Four Assistants for the Best GEO Agency. Two Named the Agency That Wrote the List They Cited.

One question, four AI surfaces, 20 domains cited and not one shared by three of them. Twice, the agency named first was the publisher of the cited ranking.

· Updated · 13 min read

We asked ChatGPT, Gemini, Google AI Overview and Google AI Mode the same question on 7 August 2026: what is the best generative engine optimization agency? Between them they cited 20 different domains. Not one of those domains was cited by three of the four.

Twice, the agency the assistant named first was the same company that published the ranking the assistant was citing.

Disclosure: EchoWi sells AI visibility measurement and GEO consultancy, so every agency named below is a competitor of ours in part of what it does. We appear in none of these answers, and this article does not rank anybody. It reports what four assistants said and which pages they drew on. Every figure was pulled with the United States and English set explicitly.


The short version

  1. 20 domains cited across four surfaces, and no domain appeared on more than two of them.
  2. Two of the four surfaces put an agency at the top of the answer while citing that agency’s own ranking page.
  3. ChatGPT was the one that named the problem, unprompted, before answering.
  4. Every surface refused to crown anyone outright, and then produced a list anyway.
  5. Three runs across an hour returned identical answers, word for word and citation for citation. That is one observation, not three, and it says something about the category.

What we ran, and what that is worth

One prompt, exactly as a buyer would type it: what is the best generative engine optimization agency? Four surfaces, read through their real interfaces rather than a provider API. United States, English, both set explicitly rather than inherited.

We ran it three times: once, again a few minutes later, and again about an hour after that. All three returned the same answers. Not similar, identical: the same agencies in the same order, and the same cited URLs in the same order, on AI Overview, AI Mode and Gemini alike.

So we have one observation, not three, and we report it as one. But the repetition is itself worth recording, because we have measured how much AI answers move between runs and found categories that need anywhere from about three runs to about thirty before the noise averages out. This category needed none. Across an hour it did not move at all.

We cannot tell from outside whether that is a genuinely deterministic answer or a cached one, and the distinction matters to an engineer and not at all to a buyer. Either way the practical consequence is the same in both directions: asking again within the hour tells you nothing new, and a sample built from rapid repeats is a sample of one wearing a disguise. If you want a distribution for this kind of question, space the runs out by days.

What one observation can support is an existence claim: this is what these four surfaces said, and these are the pages they drew on.

The citation loop

Here is the part worth the article.

SurfaceAgency named firstA source it citedSame company?
ChatGPTWebFXclutch.co, a third-party directoryNo
GeminiiPullRankyesoptimist.com, rocktherankings.comNo
AI OverviewFirst Page Sagefirstpagesage.com, its own agency rankingYes
AI ModeOptimistyesoptimist.com, its own agency rankingYes

On AI Overview and AI Mode, the first agency in the answer is the company that published the ranking page the answer is built on. The agency wrote a list, put itself at the top of it, the assistant retrieved that list, and the assistant’s answer opens with that agency’s name.

We are not saying either agency did anything wrong. Publishing a market ranking is normal, it is work, and putting yourself in it is not fraud. The observation is about the assistants: on these two surfaces, a self-published ranking was treated as a source of record, and the self-award travelled through intact.

Gemini is the interesting middle case. It cited two agencies’ own ranking pages and it named both of those agencies in its table, but neither in first position. So the citation carried, the crown did not.

This is not a rare accident. We counted the same market from the results side on the same day: thirteen of the rankings on the first page for “best geo agencies” are published by an agency that names itself first. One of the thirteen is the page AI Mode cited here. The genre is large, it is easy to find, and at least one assistant is reading from it.

ChatGPT named the problem before answering it

ChatGPT opened with this, unprompted:

“Independent rankings also vary widely, and many ‘top agency’ lists are published by agencies that rank themselves first.”

It attached a source to that sentence, then built its own list from a third-party directory and one agency blog, and named seven agencies. Notably, it cited an agency-published ranking and did not name that agency in its answer. It used the source without inheriting the source’s self-assessment.

That is the behaviour a buyer wants, and it was one surface out of four.

It also did something none of the others did: it asked for our industry, company size and budget before narrowing further, and it listed five questions to ask an agency before hiring. Whatever you think of the list it produced, the framing was the most useful of the four.

Twenty domains and almost no agreement

Across the four surfaces the answers rest on 20 distinct cited domains. The overlap is close to nothing:

Cited byNumber of domains
All four surfaces0
Three surfaces0
Two surfaces5
One surface only15

The five that appear twice are a mix: three agency-published rankings, one business directory and one software vendor’s blog. Fifteen domains were used by exactly one assistant and ignored by the other three.

Cited is not the same as retrieved, and the gap is worth knowing. ChatGPT returned nine further domains as search results that it then did not cite in its answer. We count citations here, because a citation is what a reader can click and what a publisher can be credited for. Counting everything the assistants touched would put the United States figure at 28 rather than 20, and an earlier version of this article did exactly that. The number above is citations.

This matches what we found asking the same class of question in Spanish, where 49 domains were cited across four phrasings and not one appeared in all four. Different market, different language, different question set, same shape: the assistants are not converging on an authority for this category. They are each finding a different page.

For a category with a lot of money in it, that is worth sitting with. The United States term “ai seo agency” carries 933 searches a month at a cost per click of $82.43 and a keyword difficulty of 0. That combination, a high price per click and nothing defending the term, is a market saying the customer is valuable and nobody has built authority yet. The citation spread says the same thing from the other side.

The same question in three markets

We ran the same instrument in Spain and France, with market and language declared each time, on the same day.

MarketCited domainsOn two surfacesOn threeOn all four
United States20500
France19400
Spain16400
Germany, added 7 Aug 202618220

Three markets, three languages, three separate sets of sources, and the same answer to the question that matters: no page was authoritative enough for three of the four surfaces to agree on it.

Germany, measured on 7 August 2026 and added to this table afterwards, is the first market where that failed. Two domains reached three surfaces, and both are agency websites: sem-deutschland.de, which belongs to Aufgesang, and suchhelden.de. So the honest version of the finding is narrower than the one this article published first: three of four surfaces can agree on a page, and where they do, the page is written by someone with a stake in the answer.

The markets differ in who fills the gap. In the United States and France the domains that appear on two surfaces are mostly agencies publishing their own rankings. In Spain the most cited is a business school that sells no GEO at all, and ChatGPT there warned against self-published and sponsored-press-release rankings before citing a press release. The French market has its own version of the loop, where the first agency Gemini named is the publisher of a ranking it cited.

If you are buying

Do not treat the first name in an AI answer as a verdict. On half the surfaces we checked, that name came from a page the same company wrote. Ask the assistant what it is citing. All four of these surfaces will tell you if you ask.

Ask the same question more than once, and on more than one surface. The four answers here share almost no sources, so any single one of them is a sample of one. If you are shortlisting from AI answers, take the union, not the first reply.

Ask the agency the questions ChatGPT listed, because they are the right ones: how do you measure AI visibility, can you show a client being cited, which platforms do you cover, how much of the work is off-site, and what lands in the monthly report.

If you are an agency

The honest read of this is uncomfortable in both directions.

Publishing your own ranking demonstrably works on some surfaces. Two of four opened with the publisher of the list they cited. If your goal is to be named, writing the list is a cheap way to get into the retrieval set.

And it is fragile. ChatGPT used a self-published ranking as a source and declined to repeat its verdict, while telling the reader that these lists exist and why to discount them. One assistant has already learned the pattern. Building a position on a tactic that a model can learn to discount in a single update is not a strategy, it is a trade.

The durable version is the one third-party directories are getting for free here. A directory got cited by two surfaces without publishing an opinion, because it publishes a structure a model can read: firms, categories, review counts. That is boring and it does not go stale.

We sell GEO consultancy, so read that last paragraph knowing who wrote it.

What we could not verify

  • Whether any of this is stable beyond an hour. Three identical runs tell us the answer is fixed over that span, by determinism or by caching, and nothing at all about a week or a month.
  • Whether the assistants would answer the same for a different phrasing. We ran one question. The Spanish study that used four phrasings found the source sets barely overlapped between them, so we expect this to move.
  • Why AI Overview and AI Mode chose those particular pages. Retrieval is not published. We can show the citation, not the reason for it.
  • Whether being named produces any business. Nobody has published a controlled test connecting AI citation to revenue, and a July 2026 review of 45 GEO studies found no technique with a demonstrated causal effect.
  • The quality of any agency named here. We measured citations, not work. An agency being named tells you a model retrieved a page, and nothing about whether it is good.

Common Questions About How AI Picks Agencies

Which agency do AI assistants say is the best for GEO?

They disagree, and all four decline to crown anyone before producing a list anyway. On 7 August 2026 the first name was WebFX on ChatGPT, iPullRank on Gemini, First Page Sage on AI Overview and Optimist on AI Mode. The four answers rest on 28 domains and share almost none of them.

Do AI assistants trust agency-written “best agency” lists?

Sometimes, and the consequence is visible. On two of the four surfaces we checked, the agency named first was the company that published the ranking the assistant cited. ChatGPT cited such a list and declined to repeat its verdict, warning the reader that many top-agency lists are published by agencies that rank themselves first.

How many sources does an AI answer about agencies use?

In this test, between three and eleven cited sources per surface, drawing on 20 distinct domains in total. No domain was cited by more than two of the four surfaces, and none by three. ChatGPT also retrieved nine domains it did not cite, which is why the number you get depends on whether you count citations or everything the assistant touched.

Is it worth publishing our own “best agencies” ranking?

It demonstrably gets you into the retrieval set, and on two surfaces here it put the publisher at the top of the answer. It is also the first thing an assistant learns to discount: one of the four already warns about the practice by name. Treat it as a trade with a short half-life rather than an asset.

How should I shortlist a GEO agency using AI?

Ask more than one assistant, ask each what it is citing, and take the union of the answers rather than the first name in any one of them. Then check whether the source behind a recommendation was written by the company being recommended, which took us one click per citation.

How do you measure this?

We ask the real interfaces the question a buyer would type, with market and language set explicitly, and we keep the sources each answer cites rather than only whether a brand was named. Our verified tool catalogue lists the products that do this as software, with prices read from each vendor’s own page on a stated date.

Ask an AI about this article

Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.

Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.

Written by

Maher El Ouahabi

CTO & Co-Founder at EchoWi

Builds the software that shows brands what AI is really saying about them, then what to change so the next answer is better. Twelve engines, measured before and after.

LinkedIn Maher El Ouahabi (opens in new tab)