ChatGPT Answered Two of Our Three Questions Without Citing Anything
We asked three GEO buying questions. One cited six sources, two cited nothing at all. On those, no amount of content wins you a citation.
We put three buying questions from our own category to ChatGPT. One came back with six cited sources. The other two came back with none at all. Not few. Zero. A complete, competent, well-organised answer assembled entirely from what the model already knew, with no retrieval and nothing to click.
That is the finding, and it changes what “optimising for AI” can mean on a given question. On an answer with no sources, there is no citation to win. You cannot write your way into a bibliography that does not exist.
How this was measured: Three prompts, English, United States, 6 August 2026, through the real web interfaces rather than the providers’ APIs. One run per prompt per surface, which makes each one a draw and not a rate. One surface returned an upstream error on the third prompt and is excluded from that row rather than counted as a zero. Two of the three prompts were re-run on 7 August 2026 and the results are in the correction section below, along with a number this article got wrong on the first day.
The three questions and what came back
| Prompt | ChatGPT citations | Other surfaces |
|---|---|---|
| What tools measure share of voice in AI-generated answers? | 6 | Gemini, AI Overview and AI Mode all cited sources |
| How do I track whether my brand is mentioned in ChatGPT? | 0 | Gemini 6, AI Overview 8, AI Mode 16 |
| What is generative engine optimization and does it actually work? | 0 | AI Overview errored |
The pattern is not subtle. Asked which tools do a thing, ChatGPT went and looked. Asked how to do the thing, and asked what the thing is, it answered from memory.
The uncited answers were not worse. The GEO answer ran to about eight hundred words, had a comparison table, distinguished what seems to work from what does not, and warned the reader against buying AI citations. It was the kind of answer a competent consultant gives. It just had no sources, and therefore no way for any publisher to appear in it.
Correction, and a second day of the same measurement
This article first said the cited prompt produced 28 domains. It produced six. The 28 came from a script that swept every URL out of the response record, which mixed the six sources the answer actually cited with the twenty-one results its underlying search returned. Those are different things, and telling them apart is the entire point of the section below on retrieval versus naming. We made the error this piece warns about.
The corrected figures for that prompt on ChatGPT, re-measured on 7 August 2026: six cited sources (semrush.com, ft.com, blog.hubspot.com, llmpulse.ai, get-spotlight.com and alexbirkett.com) and twenty-one results returned by the search behind the answer, most of which the answer never used. The three other surfaces cited between six and sixteen sources each; those counts came from the same sweep and we have not disaggregated them, so read them as the size of the response record rather than as citation counts.
Running the same prompts again a day later produced the second finding, and it was not what we expected.
| Prompt | 6 August | 7 August |
|---|---|---|
| What tools measure share of voice in AI-generated answers? | 6 citations | 6 citations, the same six |
| What is generative engine optimization and does it actually work? | 0 citations | 0 citations |
Both answers came back not just with the same citation behaviour but with the same text: the same table of five tools in the same order, the same section headings, the same closing paragraph. Two draws, a day apart, and nothing moved.
That matters because this site has argued repeatedly that AI answers vary between runs. Re-measured properly, that variance is real but smaller than we first reported: three runs of one Spanish AI Overview prompt cited eleven domains, five of them in every run. Both things are true, and the difference between them looks like the same split this article is about. A retrieved answer varies because the retrieval varies. A recalled answer does not vary, because nothing is being drawn.
Two runs is two runs and we are not calling that a law. It does say the question “how many runs do I need” has a different answer depending on which kind of question you are asking, which nobody in this category currently prices into their sampling.
We have since measured that properly across six categories, and the spread is larger than this section suggested: in some categories every cited source repeats in every run, and in others almost none do.
Why this is the number that matters
Every piece of advice in this category assumes retrieval. Write the quotable passage, add the schema, publish the data, earn the citation. All of it is advice about being selected from a candidate set.
If the model does not build a candidate set, none of it applies to that question. The only thing that decides whether your brand appears is whether the model already knows you, which is a property of its training corpus and not of the page you published last week.
So the first question about any prompt you care about is not “how do I rank for it”. It is “does this prompt produce citations at all”, and that is answerable in one run per surface.
Three outcomes, three different jobs:
- Cited answer, you appear. Keep the page that earned it and find out which one it was.
- Cited answer, you do not appear. This is a content and authority problem, and it is the one the whole GEO toolset is built for.
- Uncited answer. This is an entity problem. Whether the model names you depends on how well established your brand is in the corpus it was trained on, and the levers there are the slow ones: press, Wikipedia and Wikidata presence, review platforms, being written about by other people. Publishing another page does not move it this quarter.
Most tools in this market report visibility without telling you which of those three you are in. It is the difference between a fixable problem and a two-year one.
What the cited answers were built from
When the surfaces did cite, the source list was striking in a different way.
For “how do I track whether my brand is mentioned in ChatGPT”, the three surfaces that cited anything drew on 19 domains. Ranked by how many citations each received: siftly.ai, semrush.com and dageno.ai with six each, beamtrace.com with five, youtube.com with three, keyword.com with two, and one each for genrank.io, llmpulse.ai, quora.com, maxaeo.ai, frictionai.co, rankprompt.com, workduo.ai, useomnia.com, ai-semantica.com, semai.ai, nicklafferty.com, reddit.com and ahrefs.com.
Count how many of those are companies selling AI visibility tools. By our reading, thirteen of the nineteen. The answer to “how do I measure this” was assembled almost entirely from the marketing content of the people who sell the measuring.
That is not a scandal, it is a supply problem. Somebody has to write about a new category and the vendors get there first, because they are the ones with a reason to. We are one of those vendors and this article is one of those pages. The reader’s defence is not to distrust all of it, it is to check whether a claim comes with a number and a method attached.
What it means that ChatGPT looked for one question and not the others
The one prompt that triggered retrieval asked for a list of named products. The two that did not asked for an explanation and an assessment.
We have one run each, so this is a hypothesis and not a measurement: retrieval appears to be triggered by questions whose answers are lists of current, checkable entities, and skipped for questions the model considers settled definitional or advisory ground.
If that holds, it splits GEO work in two:
Product questions are winnable now. “Best X”, “X alternatives”, “tools that do Y” all pull a candidate set from the live web, and being in that set is a content and authority problem you can work on this month.
Definitional and how-to questions may not be winnable at all, on that surface, until the model is retrained. Which is an uncomfortable thing for the part of this industry that sells explainer content as GEO.
We are going to test that properly rather than leave it as a hypothesis. The test is straightforward: a set of prompts split by intent, run enough times each to produce a rate rather than a draw, reporting the share that come back with any citations at all. That number, per surface and per intent, is more useful than most of what this category currently reports, and nobody publishes it.
The uncomfortable part for us
We sell into this category, so we ran these questions knowing the answer could be unflattering, and it was.
On the two uncited questions there was no citation slot to lose, which is not a content failure and not a content opportunity either. On the cited one, the answer is exactly the problem this article describes: six sources, and the ones it chose were the sources with the longest-standing authority in the category. That is the barrier this piece is about, and it applies to us the same way it applies to the reader.
We publish that for the same reason we publish everything else. A measurement that only gets reported when it flatters you is marketing, not measurement.
What we are changing because of it
Three things, and they are all cheap.
Check for citations before writing for a prompt. One run per surface tells you whether the page you are about to write can possibly be selected.
Report the citation rate as its own metric. Not visibility, not share of voice: the share of runs for a prompt that produced any sources at all. It sets the ceiling for everything else.
Separate the two kinds of absence. Not appearing in a cited answer and not appearing in an uncited one are different diagnoses with different treatments, and rolling them into one visibility percentage hides which one you have.
What this article does not prove
- Three prompts is three prompts. One run each. This is a draw, not a rate, and we have said so at the top and here.
- One category. We asked about our own market, which may retrieve differently from yours.
- The retrieval hypothesis is untested. That product questions trigger retrieval and definitional ones do not is consistent with three observations. Three observations is not evidence of a rule.
- Model versions move. The behaviour we saw on 6 August 2026 may not be the behaviour next month, which is itself an argument for measuring continuously rather than once.
- We cannot see inside the decision. Whether the model chose not to retrieve, or retrieved and cited nothing, is invisible from the outside. We observed the absence of citations, not the absence of retrieval.
Common Questions About Uncited AI Answers
Why do some AI answers have no sources?
Because the model answered from its training rather than searching. Retrieval is a decision the system makes per question, and on questions it treats as settled, definitional or advisory, it often skips it. In our three probes, ChatGPT retrieved for a question asking which products do something and did not retrieve for a how-to question or a definition question.
Can I get cited in an answer that has no citations?
No. There is no citation slot to fill. What decides whether your brand is named in an uncited answer is whether the model already knows your brand from its training corpus, which is a matter of press, third-party coverage, reference sites and time, not of the page you publish this week.
How do I tell which kind of question I am dealing with?
Run the prompt and look at whether the answer carries sources. Do it before you commission content aimed at that prompt, because the answer decides whether the content can work at all.
Does this mean GEO does not work?
It means GEO works on the questions where retrieval happens, and that the questions where it does not are an entity-authority problem on a much longer timescale. Both are real work. They are not the same work, and a tool that reports one visibility number for both is hiding the distinction that decides what you should do next.
Who writes the sources AI cites about AI visibility?
Mostly the vendors. For one of our three prompts, thirteen of the nineteen cited domains belonged to companies selling AI visibility tools. That includes us, on other queries. The defence is not to distrust the category, it is to check whether a given claim arrives with a number and a method attached.
What should I measure instead of visibility alone?
Start with the citation rate: what share of runs for your prompt come back with any sources at all. Then split your absence into two: absent from a cited answer, which is a content and authority problem, and absent from an uncited answer, which is an entity problem. Our comparison of every tool we verified covers what each vendor publishes about how it measures, which is less than you would hope.
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.