Every LLM SEO Guide Gives the Same Eight Tips. We Measured Six of Them.
We asked two AI surfaces what LLM SEO is. They cited fourteen pages giving the same eight tips, none with a measurement. Five do not survive our register.
We asked Google’s AI Overview and Gemini the same question a buyer types: what is LLM SEO, and how do I get my site cited by ChatGPT. Between them they returned sixteen citations pointing at fourteen distinct pages. All fourteen give the same eight tips, in roughly the same order. Not one of them carries a measurement. We have seven of the eight in our register now, one of them is not measurable from outside, and five of the seven do not survive the measurement. The one thing that did separate a cited page from an uncited one turned out to separate on Google’s own surfaces and nowhere else.
Disclosure: EchoWi sells AI visibility measurement, so the advice below competes with what we do and one of the cited pages belongs to a vendor in our own catalogue. Every figure here names its population, its surface, its market and its date. The file sweeps are plain HTTP requests anyone can repeat, and the answer-layer arms publish the exact prompt, the market, the run count and the day, which is what it takes to run them again. Repeat them against us.
The short version
- Fourteen cited pages, eight tips, zero numbers. The consensus is real and it is unanimous: structure your content, allow the crawlers, add FAQ schema, get on Reddit and G2, keep it fresh, show your author, optimise for Bing, publish original data.
- “Get listed on review platforms” is the worst of the eight. Across 5,520 citation records the whole review-aggregator layer is 46 of them, 0.83%. In 31 consumer buying cells across three surfaces and two English-speaking markets, the ten biggest review platforms appear zero times.
- “Make sure you are not blocking the crawlers” solves a problem most sites do not have. Of 510 cited domains that served a
robots.txt, 79 disallow at least one AI crawler. That is 15.5%, and a slice of those rules were written by a CDN default rather than by anybody at the company. - The off-site tip is right and the platform in it is wrong. YouTube holds a citation slot in 24 of 29 category-and-market cells we measured. Reddit reaches SMB and consumer questions in 8 of 8 cells and enterprise questions in 1 of 5.
- “Structure it for machines” is the one with a controlled result behind it. Serving markdown goes with publishing an
llms.txtby 28 points, and a control on a non-AI task moves the same figure by 5.1. - The schema tip has a control, and it does not clear its own threshold. Cited pages carry a type that says what the page is 41% of the time against 31% for pages that rank and were not cited. The gap is 10 points and the threshold, written down first, was 15.
FAQPagesits on 17 pages in each group. - Nothing on the page separates a cited page from an uncited one. Rank does, and only where the answer layer shares Google’s index. Six page-level measures give tests between 0.30 and 1.34. Where the page ranks organically gives 3.47, and a frozen replication on seven questions the first arm never used gives 3.82 with the cited side ahead in 7 of 7 cells. Shuffle the cited sets between questions and 0 of 135,350 possible shuffles reach either result, so it is not a handful of big domains ranking well everywhere. Then put the same nine questions to Gemini, which this register has measured as retrieving separately, and rank stops separating: 1.07, and 4 of 8 cells. Nothing this study can measure predicts a citation on an independent retriever.
- One of the eight we have not measured, and we say which.
Where the eight tips come from
The question was asked once, on 25 August 2026, from the United States in English, of AI Overview and Gemini. AI Overview returned ten citations and Gemini six. Two pages appear on both surfaces, so the sixteen citations point at fourteen distinct URLs, and the count below is of pages and not of citations. The pages are guides from agencies, consultants and vendors, plus one YouTube video in first position.
Flattened, the advice is:
| # | The tip | In our register? |
|---|---|---|
| 1 | Allow GPTBot and the other crawlers in robots.txt | Yes |
| 2 | Answer first: put the direct answer under the heading | Yes, indirectly |
| 3 | Add FAQPage schema so machines can parse the questions | Yes, with a control |
| 4 | Earn mentions on Reddit, G2, Wikipedia, directories | Yes |
| 5 | Keep content fresh and update timestamps | Yes, with a control, and it fails |
| 6 | Show author credentials | Yes, with a control, and it does not separate |
| 7 | Optimise for Bing, since ChatGPT leans on that index | Not measured |
| 8 | Publish original data and first-hand insight | Yes |
That is six we can settle, one we can speak to, and one we cannot. The one we cannot is stated as such rather than repeated.
Tip four is the one to drop: the review layer barely exists
The advice to get listed on G2, Capterra and Trustpilot appears on both surfaces, cited to different pages. It is the single most confidently repeated tip in the set, and it is the one our data contradicts hardest.
Across every citation this register has recorded, 5,520 records over 1,801 distinct domains, the ten biggest review and aggregator platforms account for 46 records, 0.83%. G2 is 18 of those, Gartner 16, Capterra 9, GetApp 3. Software Advice, TrustRadius, Trustpilot, Product Hunt, SourceForge and Clutch are on zero. In the same body of readings, YouTube is 330 and Reddit is 195.
That number could be an artefact of what we happen to ask, since most of those readings are B2B software purchases. So we tested it where the layer should be strongest: eight consumer buying questions, chosen before measuring for being the ground where a complaints platform lives, car and home insurance, an electricity supplier, a mobile plan, a VPN, a mattress, a budget airline and a meal-kit box.
| Surface and market | Cells | Any of the ten platforms cited |
|---|---|---|
| AI Overview, United States | 8 | 0 |
| AI Overview, United Kingdom | 7 | 0 |
| AI Mode, United States | 8 | 0 |
| Gemini, United States | 8 | 0 |
| AI Overview, Denmark, in Danish | 7 | 1 |
The Danish cell is the entire exception, and it is one cell of seven and one run of three. Trustpilot is a Danish company, so the layer exists where the platform is a household name in the household’s own language, and nowhere else we looked.
One reading is worth keeping from the UK sweep: on the home-insurance question the answer named Trustpilot in its text, in all three runs, and cited no page of theirs in any of them. Being named and being a source are different outcomes. A tool that reports the first as the second tells a vendor it is winning a layer it is not in.
Tip one solves a problem about one site in six has
“Check your robots.txt and make sure you are not blocking GPTBot” is sound advice in the sense that the failure it describes is real and fatal. It is also advice for a condition most of the sites getting cited do not have.
We took the 547 domains this register has recorded as cited three or more times, and read each one’s robots.txt against fourteen named AI crawlers. Of the 510 that served a file, 79 disallow at least one of the ten agents the earlier study checked, 15.5%. Googlebot is allowed on 509 of the 510, which is the control: a parser reading the rules wrong would show Googlebot blocked at something like the AI rate, and it does not.
Two things sit underneath that number and neither is in any of the fourteen guides.
The first is that naming a crawler is far more common than blocking one. Of 1,132 domains in the largest stratum, 306 name an AI crawler rule and only 104 disallow one. Two out of three sites that have written a policy have written a permission.
The second is that a good share of the refusals are not decisions. Reading the rules that decide each verdict, a slice of the blocking domains have theirs inside a # BEGIN Cloudflare Managed content block, which is a hosting default and not a sentence anybody at the company wrote. In the sibling convention the effect is stronger still: of the sites that refuse AI training via a Content Signal, most did not type the refusal.
So the honest version of tip one is narrower and more useful: check the file, and then check who wrote what is in it, because your CDN may have answered the question for you.
The off-site tip is right, and the platforms in it are the wrong ones
Every one of the fourteen pages says to build presence off your own domain, and every one of them illustrates it with Reddit, G2 or Wikipedia. The instinct is correct. The examples are ranked wrong.
YouTube is the largest single slot we have measured. Across 29 category-and-market cells, YouTube holds a citation slot in 24. The next platform is Reddit at seven, and after that it falls to three. That is a count of the domain, and we have narrowed it ourselves at page level: in one cell the five YouTube links are four videos from three channels, so the slot is smaller than the domain count suggests. In the very answer that produced this article, the first source AI Overview cited was a YouTube video.
Reddit works, and it works for a specific buyer. Across the categories we tested, questions from consumers and small businesses reached Reddit in 8 of 8 cells. Enterprise questions reached it in 1 of 5. If you sell to procurement, the Reddit advice is being given to you by somebody who measured it on a different buyer.
And the slot is not the same thing as a durable position. In one market and category we read the video slot at depth: YouTube held it in every run of a four-run AI Overview reading, and in half the runs of a six-run AI Mode reading of the same question on the same day. Broad presence, intermittent hold.
The schema tip has a control, and it does not clear its own bar
Adding FAQPage schema is the third tip on every one of the fourteen pages, and
it is the one where the advice and the evidence are furthest apart, because it
is the only tip in the set that somebody has actually run a control on.
We put the same instrument to two groups: the pages an AI answer cited, and the pages that ranked for the same nine questions and were not cited. The direction and the bar were written down before the first fetch, and the bar was 15 points.
| Cited pages | Ranked, not cited | |
|---|---|---|
| Any JSON-LD at all | 93% | 77% |
| A type that says what the page is | 41% | 31% |
| Nothing at all | 7% | 14% |
The gap on the types that say what a page is is 10 points, and the bar was
15. All three rows lean the same way, and none of them leans far enough to
call markup the thing that separates the two groups. FAQPage itself sits on 17
pages in each group, which is as flat as a comparison gets.
The honest version of the tip is therefore narrower than any of the fourteen pages give it: fewer than half the pages an engine cites declare what they are, so the schema is not the price of admission, and 3 of the 46 cited pages carry no structured data at all while running to 29,724, 12,429 and 6,737 characters of text. The full reading, including the bot-challenge difference that turned up alongside it, is in what cited pages declare about themselves.
The freshness tip is true about a field and not about the writing
Keep it fresh is the most repeated of the eight. It is also the only one whose answer flips depending on which date you call the date.
We read all 71 URLs an AI answer returned as a source in our own measurements, the same population the schema sweep used, on the same day. 45 answered with a body a crawler can read. 37 of those 45 carry a machine-readable date, so the negation dies immediately: it is not the case that the pages winning citations have nothing to update.
Then the number splits in two.
| Which date | Pages | Median age |
|---|---|---|
| Newest date of any kind | 37 of 45 | 49 days |
| Publication date | 36 of 37 | 419 days |
| Modified stamp | 34 of 37 | 46 days |
The pages an AI answer cites were written 419 days ago at the median and stamped 46 days ago. 33 carry both fields, and on those the stamp sits 426 days after the publication at the median. 20 of the 33 were written more than a year ago, and 13 of those 20 carry a stamp from the last ninety days. The oldest page in the set was published in 2013, carries a stamp from last month, and was cited today.
So the tip is supported in its literal form and it is advice about a field. A
page written fourteen months ago with a stamp from last month is what wins, and
nothing here can tell you whether that page was genuinely rewritten or simply
re-stamped. dateModified is changed without touching a word.
The design was frozen before the first request with four predictions and their thresholds, and all four held, including the instrument control: the schema sweep read 46 of the same 71 URLs and this one read 45, two points apart against a threshold of fifteen. What the frozen design did not anticipate is that the statistic’s unit decides the answer. It named the newest date, which gives 49 days and flatters the tip. The publication date gives 419, which would have failed the same prediction. Both are defensible and we published the one the design named, and then the one it did not, rather than picking.
The rows are in our measurement register.
The control arm: cited and uncited pages have the same dates
The section above is one arm. It is selected on the outcome, so it can say the pages that won carry recent stamps and it cannot say that carrying one is why they won. The arm that answers that is the one we had named with its cost, and this is it.
For each of the same nine prompts, in the same nine markets, we asked Google for the organic results and kept every page whose registrable domain was not already in the cited set for that cell. That is 86 pages, read the same day with the same instrument.
42 of the 142 organic results were already cited. So the sets overlap by about 30 per cent, which is worth knowing on its own: ranking and being cited are neither the same thing nor unrelated.
| Cited | Ranked, not cited | |
|---|---|---|
| Read | 45 of 71 | 62 of 86 |
| Carry a machine-readable date | 82% | 71% |
| Median publication age | 419 days | 761 days |
| Median modified stamp | 46 days | 71 days |
Read the medians and the tip looks vindicated: the cited pages were written 342 days more recently and stamped 25 days more recently. Read the distributions and none of it is there. On a rank test the publication age gives z = 0.61 and the modified stamp z = 0.30. The share carrying any date at all differs by 11 points at z = 1.34. On every axis the two groups are indistinguishable.
The medians move because these distributions have long tails and the samples are small. One arm has a page from 2013 in it, and a median of forty-odd values slides a long way on that.
So the answer to the most repeated of the eight tips is that the pages an AI cites and the pages it passed over, for the same question on the same day, do not differ in their dates in a way this test can detect. That also explains the one-arm result: 82 per cent carrying a date and a 46-day median stamp is not what a cited page looks like, it is what a page on this kind of query looks like.
Two things this does not say. It does not say dates are irrelevant, only that they do not separate these two groups at this sample size, and a real effect smaller than the noise would look exactly like this. And the two arms are not the same kind of page: 18 of the 62 uncited pages we could read carry no date anywhere, and they are mostly retailer and brand category pages, which are structurally dateless. That is declared rather than corrected for.
The rows for both arms are in our measurement register.
Tip eight is the one with a controlled result behind it
“Publish original data” and “structure your content so machines can extract it” are the two tips the fourteen pages agree on most and evidence least. The second one we can now speak to with a control, which is rarer than it should be in this field.
Of 466 cited domains that answered a negotiated homepage request, 45 serve markdown, 9.7%. Those that do publish an llms.txt 67% of the time against 39% for those that do not, a gap of 28 points.
The obvious objection is that both are marking the same underlying thing, a modern technical site that does optional work, rather than anything about AI. So we asked the same 547 domains for a task that is optional, standardised and has nothing to do with AI: /.well-known/security.txt. The design and its thresholds were written and committed before the first request.
90 of 479 domains that answered publish one, 18.8 per cent. And it predicts nothing: the llms.txt gap for security.txt publishers is 5.1 points with z = 0.87, against 28 points for markdown. Naming an AI crawler rule moves 4.3 points, and blocking one moves 3.7. All three land in the noise.
Doing optional standards work does not predict any AI behaviour. Doing an AI-facing one does. That is as close to “this tip is about AI and not about being tidy” as an observational sweep gets.
The author tip, measured with a control, and the unit is the whole story
“Show your author credentials” is the sixth tip and it was the last of the eight this register could answer. It is a good test of our own discipline, because the answer depends entirely on what you count, and we wrote down what we would count before we fetched anything.
Three rules, frozen with the predictions: an author that is only an
Organization is not a person; a Person carrying nothing but a name is a byline
and not credentials, and the tip says credentials; a bare string author has no
type at all, so it cannot be a person and gets its own column. A
<meta name="author"> tag is a separate signal and no total adds it in. The
decisive statistic, named in advance, is the share of readable pages whose
JSON-LD declares an author of type Person.
Both arms read on 26 August 2026 with the same instrument, the same four states and the same retry rule as the freshness sweep.
| Cited | Ranked, not cited | Test | |
|---|---|---|---|
| Pages asked | 71 | 83 | |
| Readable | 45 | 61 | |
| Declares a Person author | 38% | 28% | z = 1.08 |
| Of those, with credentials | 11 of 17 | 6 of 17 | z = 1.71 |
| Declares any author at all | 80% | 54% | z = 2.77 |
The decisive statistic does not separate. 38 against 28 with z = 1.08, which is the sixth page-level axis in a row to land inside the noise and sits neatly in the 0.30 to 1.34 band the other five occupy.
And then the row underneath it does. “Declares any author at all” separates at z = 2.77, which would have been the first page-level measure to clear the bar if it had been the one we chose. It was not the one we chose, it is a post-hoc observation, and the only honest way to promote it is to freeze it as a prediction on a sample where it was not found. So it is written here as an observation and nowhere else.
There is also a good reason to doubt it, and it is visible in the same table. The cited arm returned 20 shells against the control’s 7, and the control arm carries the category and store pages the freshness sweep already counted as structurally undated. Those pages have no author because of what they are, not because of anything their publisher decided. The gap in “any author” is very probably that composition difference wearing an author’s clothes.
Which makes the null sharper rather than weaker. Even with the cited arm tilted toward editorial pages, the statistic that describes what the tip actually asks for does not separate the two groups.
The credentials row is the one that speaks to the tip’s own wording, and the
denominators are seventeen a side, so it is an observation and not a result:
among pages that name a person, 11 of 17 cited pages attach something like a job
title, an employer or a sameAs, against 6 of 17 uncited ones. Our prediction
said fewer than half and the retirement threshold said two thirds; 65 per cent
landed between them, in a band the frozen design did not name. That is a defect
in the design and it is written down as one rather than resolved now that the
number is visible.
Six things we can measure about a page, and none of them separates
By this point the register has four studies on the same two arms: the same pages an AI answer cited, and the same pages that ranked for those questions and were not cited. Between them they measure six things about a page. Here they are on one scale.
| What we measured | Cited | Ranked, not cited | Test |
|---|---|---|---|
| Declares a citable-unit schema type | 41% | 31% | z = 1.14 |
| Carries any machine-readable date | 82% | 71% | z = 1.34 |
| Median publication age | 419 days | 761 days | z = 0.61 |
| Median modified stamp | 46 days | 71 days | z = 0.30 |
| Median readable length | 28,605 chars | 22,803 chars | z = 1.04 |
| Declares an author of type Person | 38% | 28% | z = 1.08 |
Every row leans the same way. Not one of them clears the bar. The cited pages carry more schema, carry dates more often, were written more recently, were stamped more recently, are 25 per cent longer and name a person as author more often, and on every axis the two distributions overlap enough that the difference is inside the noise.
That consistency is worth naming rather than dressing up. Six measures all leaning the same direction is not what pure chance looks like, and it is also not six independent coin flips: longer pages carry more schema, more dates and more bylines, so the six are correlated and “all six lean the same way” is far weaker evidence than six independent tests would be. We are not going to do the arithmetic that pretends otherwise.
What we will say is the honest reading of the pattern. There is probably a small real effect on each of these axes, and not one of them is large enough to be the thing that decides whether a page gets cited. The advice in all fourteen pages we started from is a list of these axes. If the effect of each is smaller than the run-to-run noise of the measurement, then following the whole list is not what separates a cited page from an uncited one, and the explanation is somewhere we did not look: off the page.
The length axis was frozen before it was computed, with the synthesis condition written in advance so it could not be chosen afterwards: if all three studies failed to separate, they get published as one finding; if any one separated, that one was the headline and there was no synthesis. All three failed.
The one thing that does separate is not on the page
If nothing we can measure about a page separates the two groups, the next place to look is off it. The most obvious off-page candidate is the one we already had the data for: where the page ranks.
Of the 142 organic results across those nine queries, 42 have their registrable domain in the cited set for that same question and 100 do not. So the comparison runs inside one SERP, with no new arm to build.
| Cited domain | Not cited | |
|---|---|---|
| Median organic rank | 8 | 13 |
| In the top ten | 69% | 39% |
Rank test z = 3.47. Not 1.04, not 1.34. 3.47.
And because 142 results come from only nine independent queries, treating them as 142 independent observations overstates the power. The frozen design said so and named a per-cell sign test as the tiebreaker: in how many of the nine cells is the median rank of the cited domains better than the uncited ones? Eight of nine, and the ninth is a tie at 11 against 11, not a reversal. The clustered test agrees with the pooled one.
So the shape of the whole series is this. Five things the advice tells you to change about your page, and not one of them separates a cited page from an uncited one. One thing the advice barely mentions, and it separates at three times the strength.
Three things that stops short of, and all three matter.
It is not causal. Google’s AI Overview and AI Mode are assembled from the same index that produces those organic results, so an answer layer drawing from the top of its own SERP is partly mechanical. This is a description of where citations come from, not proof that moving up causes them.
Rank is downstream of everything, including the five page measures. So this does not say that the page does not matter. It says that in this sample, once you know where a page ranks, its schema and its dates and its length tell you nothing more, and it is entirely possible that our five measures are simply poor proxies for whatever earns the rank in the first place.
And the matching here is by registrable domain, not by URL, so cited means this domain was cited for this question rather than this exact page was. That is looser than the six page-level axes. We went and tightened it, and the section below reports what happened.
What survives all three caveats is still worth the buyer’s time: the corpus of advice that sent us here spends eight tips on the page and almost none on the position, and the position is the only column in our table that moves.
The rank rows are in our measurement register, organic rank.
The rank result replicates on a set of questions it has never seen
A result that strong from nine questions read on one day is exactly the kind of thing that does not survive being asked again, so we froze a replication before running it: a set of questions the original never used, the thresholds written down first, and the same two statistics so the two arms read on one scale.
The set comes from a rule and not from a shortlist. Every cell in the register where AI Overview answered all three runs and left four or more sources standing, minus the ones the first arm already used. That leaves business banking, CRM, ecommerce, email marketing, hosting, password managers and project management.
Three things differ from the first arm, and two of them make the test harder. One market instead of four, so this replicates the question axis and not the market axis. Sixteen days between the two halves: the cited sets were measured on 10 August 2026 and the SERP was read on 26 August 2026, against the much shorter gap in the original. And a stricter definition of cited: here it means the domain was cited in every answered run, not that it turned up as a source once.
The replication returns 115 organic results across 7 questions, of which 40 have their registrable domain in that question’s cited set and 75 do not.
| First arm | Replication | |
|---|---|---|
| Organic results | 142 | 115 |
| Cited domains among them | 42 | 40 |
| Median organic rank, cited | 8 | 8 |
| Median organic rank, uncited | 13 | 14 |
| In the top ten, cited | 69% | 70% |
| In the top ten, uncited | 39% | 32% |
| Rank test z | 3.47 | 3.82 |
| Cells where cited ranks better | 8 of 9 | 7 of 7 |
The frozen thresholds were z of 2.0 or more to replicate and under 1.5 to refute, a median gap of at least 3 positions, and the cited side better in 5 or more of the 7 cells. All three clear, and the sign test is unanimous, with no tie and no reversal.
What that buys is not a bigger number. It is that the one axis in this series that separates does it across two disjoint sets of questions, sixteen days apart, under a stricter definition of being cited, with the direction written down before the data existed.
The caveats do not move. It is still not causal, rank is still downstream of the five page measures, and cited still means this domain was cited for this question rather than this exact page. What has gone is the possibility that the first result was a one-day, nine-question accident.
Both arms are in our measurement register, organic rank.
The control that could have killed it: give each question somebody else’s cited set
Everything above varies the context. More questions, another market, a stricter definition of cited, sixteen days between the halves. None of it varies the thing being claimed, and the thing being claimed is that the cited domains of that same question rank better than the rest of that question’s results.
There is an obvious alternative that all of it survives: big domains rank well and get cited everywhere, so any set of well-known domains would look good against any SERP. If that is what we are measuring, the relationship is not between a question and its answer, it is a property of a handful of large sites.
The control is to break the pairing. Give each question the cited set of a different question and recompute. If the effect is per-question, it has to die. It costs no calls at all, because it is a shuffle of rows already in the register, so every possible shuffle that leaves no question with its own set is enumerated rather than sampled: 133,496 of them in the first arm and 1,854 in the second.
| First arm | Replication | |
|---|---|---|
| Real rank test z | 3.47 | 3.82 |
| Median z across every shuffle | 0.51 | 0.88 |
| Highest z any shuffle reached | 2.11 | 1.58 |
| Shuffles that reached the real z | 0 of 133,496 | 0 of 1,854 |
Not one shuffle out of 135,350 gets near either result. The real value is not merely above the 95th percentile of the shuffled distribution, it is above the maximum of it, in both arms independently.
One number in that test does not go to zero and it is worth saying why. In the replication, the median shuffled gap between the two medians is still 3 positions, where in the first arm it is 0. That is the confounder showing its actual size: a few domains do appear in the cited set of several questions, so a borrowed set still catches some of them. The rank test prices that in and returns 0.88 against a real 3.82, which is what the control is for, and it is a reason to read the z and the sign test rather than the gap on its own.
What this does not fix is the caveat the study started with. Google’s answer layer is built from the same index that produces these results, so it can still be drawing from the top of its own SERP. The shuffle rules out the cheapest alternative explanation, not the mechanism.
Is it the rank, or is it page one? Both verdicts were written down first
“Move from eight to three” and “get onto page one” are the same sentence if the relationship is a step and different advice if it is a gradient, and neither the median rank nor the top-ten share tells them apart.
The obvious test is the wrong one, and it is worth saying why before running anything. Rerunning the rank test on the top ten alone conditions on the very variable being tested: it compresses the range and shrinks the statistic by construction, so a small number there would mean nothing. What does not have that defect is a rate. Split the results into rank bands and ask what share of each band is a cited domain. The bands were fixed before the split was seen, and so were both verdicts: a gradient means the share falls at every step in both arms, a step means the three page-one bands sit within 10 points of each other while the second page sits at least 20 below.
| Rank band | First arm | Replication |
|---|---|---|
| 1 to 3 | 57% (8 of 14) | 70% (7 of 10) |
| 4 to 6 | 38% (8 of 21) | 40% (6 of 15) |
| 7 to 10 | 39% (13 of 33) | 56% (15 of 27) |
| 11 to 20 | 18% (13 of 74) | 19% (12 of 63) |
Neither verdict is met, and that is published rather than softened. The share does not fall at every step, because 4 to 6 sits below 7 to 10 in both arms. And the three page-one bands are 19 points apart in the first arm and 30 in the replication, well outside the 10 that the step definition allowed.
What the numbers do say is clearer than either label. Pooling page one against the second page: 43% against 18% in the first arm and 54% against 19% in the replication, with rank tests of 3.27 and 3.90 and Wilson intervals that do not touch, 32 to 54 against 11 to 28 and 41 to 67 against 11 to 30. Inside page one, comparing the top three against ranks 4 to 10 gives 1.23 and 1.14, and all six page-one intervals overlap.
So the boundary is where the difference lives, and within page one this sample cannot tell positions apart. The strict step definition failed on its tolerance and not on its idea: 10 points was too tight for bands of 10 to 33 results, which is a fact about the threshold I wrote, not about the web. The honest reading is that getting onto page one is worth something measurable here and moving from eight to three is not something this data can support.
Both arms and all eight bands are in our measurement register, organic rank.
The weakest part of the rank result, tightened: match the page, not the site
The rank finding has one caveat the other six measures do not: it pairs by registrable domain. A result counted as cited if that site was cited for that question, not if that page was. Every page-level axis in the table above uses the page as its unit. This one did not, which made the strongest number in the study also the loosest in its unit.
So we matched by exact URL. The design, the normalisation rule and all four
predictions with their middle bands were committed before anything was
computed. The rule is deliberately narrow: lowercase the host, drop www., the
scheme, the query, the fragment and a trailing slash, and follow no redirects,
because collapsing near-matches after seeing which ones fail is choosing the
pairing.
One SERP reading of the same nine questions, both pairings computed on it, so the only thing that differs between the two rows is the unit.
| Pairing | Matched | Median rank, matched | Median rank, rest | Rank test | Cells where matched is better |
|---|---|---|---|---|---|
| Exact page | 31 of 127 | 6 | 12 | z = 3.90 | 8 of 9 |
| Registrable domain | 36 of 127 | 6 | 12 | z = 3.75 | 7 of 9 |
Tightening the unit made the result stronger, not weaker. The exact-page pairing gives a larger test statistic and one more cell going the same way, on a strictly smaller matched set.
And the prediction we got wrong is the informative one. We predicted the page pairing would lose more than half its matches, because a site that gets cited can easily rank with some other page. It loses 5 of 36. So the domain pairing was barely looser than the page pairing in practice, which is why the two rows sit on top of each other, and the word “looser” that the caveat above carries is worth about five results in nine questions.
The internal control is the row underneath. The domain pairing recomputed on this reading gives 3.75 against the 3.47 the study publishes on its own reading, which is what lets the two pairings be compared at all: had it landed far away, today’s SERP would not be the one that produced the study and neither number would mean anything next to the other.
What this still does not settle is unchanged and worth repeating, because it is the part a reader should carry away. It is not causal; Google’s answer layers draw on the index that produces these results, so a layer serving the head of its own SERP is part of the mechanism rather than a rival explanation. And the cited sets were read on earlier dates than this SERP, so a page cited then and replaced since counts here as not cited, which pushes the page pairing down rather than up.
The control that changes the index instead of the pairing
Every control so far has changed something around the rank result. More questions. Another market. A stricter definition of cited. Somebody else’s cited set shuffled in. None of them changed the thing the result is most likely to be an artifact of, which is written into the caveat above: AI Overview and AI Mode are built from the same index that produces the organic results we paired against. A layer serving itself from the head of its own results page would produce exactly this number and mean nothing by it.
There is one surface where that objection does not apply. This register has Gemini measured as retrieving separately: a median of 37 per cent of an AI Overview cited set sits inside Gemini’s, against 94 per cent for AI Mode. So we asked the same nine questions of Gemini, in the same four markets, on the same day as the SERP, and labelled the same 142 organic results a second time.
The design, the statistic, four thresholds and the reading order were frozen and committed before the first call. The reading order matters more than usual here, because one of the four thresholds can make the other three meaningless.
| Answer layer | Cited domains among the results | Median rank, cited | Median rank, rest | z | Cells where cited is better |
|---|---|---|---|---|---|
| Shares the index with the SERP | 42 of 142 | 8 | 13 | 3.47 | 8 of 9 |
| Retrieves separately | 27 of 142 | 7 | 11 | 1.07 | 4 of 8 |
The overlap control is read first. If Gemini had cited what AI Overview cited, this would be Google measured twice and the row below it would say nothing. It did not: the two cited sets share a median Jaccard of 0.15, the highest single cell is 0.43, and one cell has nothing in common at all. That is a second retriever.
Which means the second row can be read, and the second row does not replicate. The rank test gives 1.07, against 3.47 and 3.82 on the two Google surfaces, and the sign test comes out 4 of 8, a coin flip. The medians still lean the right way, 7 against 11, and the sign test governs when the two disagree, which was also written down first. Medians of forty-odd values slide a long way on a handful of long-tail results; this study has now watched that happen twice.
So the caveat stops being a caveat and becomes the finding. Organic rank predicts citation on the surfaces built from the same index as the results page, and does not predict it on the one surface that retrieves separately. Put beside the six page-level axes, that leaves something uncomfortable and worth saying plainly: nothing this study knows how to measure predicts a citation on an independent retriever.
Two honest deductions from the arm. One cell of the nine, the French headphones question, has no Gemini-cited domain anywhere in its SERP, so the sign test runs on eight pairs and not nine. And two cells answered two runs of the three requested, which is recorded rather than smoothed over.
That French cell is the most quotable thing here and nobody predicted it. Asked
in French, from France, for the best noise-cancelling headphones, Gemini answers
out of journaldemontreal.com, pcmag.com, cnet.com, stuff.tv,
whathifi.com and crutchfield.com. The French results page for the identical
string on the identical day is lesnumeriques.com, frandroid.com,
01net.com, fnac.com, boulanger.com. Not one domain in common, and not
because Gemini was quiet: it cited six domains in that cell, and across the
study 27 of 142 ranked results sit in one of its sets against 42 for AI
Overview.
If you sell in France and you are watching your position on Google, that cell is the whole argument for watching more than one surface.
What a buyer does with this
If you are being sold an LLM SEO engagement, the eight tips are what you are buying. Three questions separate the people who measured from the people who read the same fourteen pages you can read for free.
- “Which of these have you measured, and on what population?” A guide is a hypothesis. Ask for the denominator, the surface, the market and the date. If the answer is a best-practice list, you are paying for a summary of a search result.
- “Show me the platform ranking, not the platform list.” Reddit, G2 and YouTube are not interchangeable, they are not equally sized, and which one matters depends on whether your buyer is a consumer or a procurement team.
- “Is that a mention or a citation?” They are different outcomes and the difference is measurable. Our own AI visibility platform reports them apart for exactly that reason, because a dashboard that merges them will show you winning a layer you are not in.
What this does not show
- Fourteen pages, one question, one day. The consensus is what two surfaces cited on 25 August 2026 in the United States in English. Another question, market or day would cite different pages, and quite possibly the same eight tips.
- The rank finding is now bounded, and the bound is the interesting part. It holds on the two surfaces built from the same index as the results page and does not hold on the one that retrieves separately. Tightening the pairing from the site to the exact page made it stronger; changing the index made it disappear.
- Nothing causal. Every figure here is observational. Domains that serve markdown also publish
llms.txtmore often; neither causes the other, and both may follow from something we did not measure. - Populations selected on the outcome. The crawler,
llms.txt, markdown andsecurity.txtsweeps all run on domains this register has already recorded as cited. That is a useful population for asking what cited sites look like, and it is not a sample of the web. - Adoption is not use. Publishing a file is not the same as anything reading it, and we have observed no fetches of a per-page markdown twin in a week of edge logs.
- One tip we did not test, and we cannot test it from here. The Bing tip stands on ChatGPT’s retrieval, and ChatGPT is the one surface whose sources our sweeps cannot read. It is absent above rather than dismissed, it is worth testing by anyone who can see inside that retrieval, and repeating it here without a number would have made this the fifteenth page.
Common Questions About LLM SEO
What is LLM SEO?
LLM SEO is the practice of shaping a site so that large language models retrieve it, understand it and cite it in their answers. It is the same activity as generative engine optimization and answer engine optimization; the three names describe one job, which is being the source an assistant uses rather than a link on a page of results.
Does getting listed on G2 or Trustpilot help AI visibility?
Not on the evidence we have. Across 5,520 citation records the ten biggest review platforms are 46 of them, 0.83%, and in 31 consumer buying cells across three surfaces and two English-speaking markets they appear zero times. The single exception we found was in Denmark, in Danish, where Trustpilot is a domestic brand, and it was one cell of seven at one run of three.
Is my robots.txt blocking ChatGPT?
Probably not. Of 510 cited domains that served a robots.txt, 79 disallow at least one AI crawler, 15.5%. The more useful check is who wrote the rule: some refusals sit inside a CDN’s managed block and were never typed by the site owner. Read the file, find the group that decides each agent, and see whether the line is yours.
Should I publish an llms.txt?
It is cheap and roughly half the cited web has done it: adoption is 41.6%, 43.8% and 44.0% across three disjoint strata of cited domains. What it does not do is guarantee a reader. Treat it as a low-cost convention rather than a lever, and ask any vendor who sells it to you for a fetch log.
What is the difference between being mentioned and being cited?
A mention is your name appearing in the answer text. A citation is a page of yours appearing in the sources. We have watched one answer name a review platform in all three runs while citing none of its pages, which is exactly the case where a tool that merges the two reports a win that did not happen.
How do I measure any of this rather than assume it?
Pick the questions your buyers actually ask, run them repeatedly against each surface separately, and record which pages come back, not just whether your name did. One run is a draw and not a rate. The register behind this article is published in full, including the readings that went against us.
Ask an AI about this article
Opens your assistant with this page already loaded, so you can check the numbers, argue with the method or ask what it means for you.
- ChatGPT (opens in new tab. the question is pre-filled, press enter to send it)
- Claude (opens in new tab. the question is pre-filled, press enter to send it)
- Perplexity (opens in new tab)
- Google AI Mode (opens in new tab)
Perplexity and Google answer straight away. ChatGPT and Claude fill the box and wait for you to press enter, which is their behaviour and not something we can set.